Method and device for detecting and correcting semantic error in wireless communication system

The method and apparatus address semantic error detection and correction in wireless communication systems using attention maps, improving reliability and efficiency in advanced communication technologies.

WO2026049076A1PCT designated stage Publication Date: 2026-03-05LG ELECTRONICS INC

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-27
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing wireless communication systems face challenges in efficiently detecting and correcting semantic errors, particularly in advanced communication technologies like enhanced mobile broadband (eMBB) and massive machine type communications (mMTC), which require reliable and latency-sensitive services.

Method used

A method and apparatus for detecting and correcting semantic errors in wireless communication systems using attention information and prior information, generating an attention map based on input data to facilitate efficient semantic communication.

Benefits of technology

Enables effective semantic communication by accurately detecting and correcting errors, enhancing the reliability and efficiency of wireless communication systems, particularly in advanced technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024012757_05032026_PF_FP_ABST
    Figure KR2024012757_05032026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure is to detect and correct a semantic error in a wireless communication system. A method performed by a first device may comprise the steps of: establishing a connection with a second device; receiving a first message requesting capability information from the second device; transmitting a second message including the capability information to the second device; receiving configuration information for communication from the second device; and performing semantic communication supporting at least one task on the basis of the configuration information.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for detecting and correcting semantic errors in wireless communication systems

[0001] The following description relates to a wireless communication system, and to a device and method for detecting and correcting semantic errors in a wireless communication system.

[0002] Wireless access systems are widely deployed to provide various types of communication services, such as voice and data. Typically, wireless access systems are multiple access systems that support communications with multiple users by sharing available system resources (e.g., bandwidth, transmission power). Examples of multiple access systems include code division multiple access (CDMA), frequency division multiple access (FDMA), time division multiple access (TDMA), orthogonal frequency division multiple access (OFDMA), and single-carrier frequency division multiple access (SC-FDMA).

[0003] In particular, as numerous communication devices demand greater communication capacity, enhanced mobile broadband (eMBB) communication technologies are being proposed, surpassing existing radio access technology (RAT). Furthermore, massive machine type communications (mMTC), which connects multiple devices and objects to provide diverse services anytime and anywhere, as well as communication systems that consider reliability and latency-sensitive services / user equipment (UE), are being proposed. Various technological configurations are being proposed for these purposes.

[0004] The present disclosure relates to a method and device for effectively performing semantic communication in a wireless communication system.

[0005] The present disclosure relates to a method and apparatus for efficiently detecting semantic errors in a wireless communication system.

[0006] The present disclosure relates to a method and apparatus for efficiently correcting semantic errors in a wireless communication system.

[0007] The present disclosure relates to a method and apparatus for detecting and correcting semantic errors based on attention information in a wireless communication system.

[0008] The present disclosure relates to a method and apparatus for detecting and correcting semantic errors by utilizing prior information in a wireless communication system.

[0009] The present disclosure relates to a method and apparatus for generating prior information for input data for detecting semantic errors in a wireless communication system.

[0010] The technical objectives to be achieved in the present disclosure are not limited to those mentioned above, and other technical tasks not mentioned can be considered by a person having ordinary skill in the technical field to which the technical configuration of the present disclosure is applied from the embodiments of the present disclosure described below.

[0011] As an example of the present disclosure, a method performed by a first device in a wireless communication system may include the steps of establishing a connection with a second device, receiving a first message requesting capability information from the second device, transmitting a second message including the capability information to the second device, receiving configuration information for communication from the second device, and performing semantic communication supporting a plurality of tasks based on the configuration information. The configuration information may include information for performing semantic error detection and correction using an attention map generated based on prior information of input data.

[0012] As an example of the present disclosure, a method performed by a second device in a wireless communication system may include the steps of establishing a connection with a first device, transmitting a first message requesting capability information to the first device, receiving a second message including the capability information from the first device, transmitting configuration information for communication to the first device, and performing semantic communication supporting a plurality of tasks based on the configuration information. The configuration information may include information for performing semantic error detection and correction using an attention map generated based on prior information of input data.

[0013] As an example of the present disclosure, in a wireless communication system, a first device includes a transceiver and a processor connected to the transceiver, wherein the processor can establish a connection with a second device, receive a first message requesting capability information from the second device, transmit a second message including the capability information to the second device, receive configuration information for communication from the second device, and perform semantic communication supporting a plurality of tasks based on the configuration information. The configuration information can include information for performing semantic error detection and correction using an attention map generated based on prior information of input data.

[0014] As an example of the present disclosure, a communication device includes at least one processor, and at least one computer memory connected to the at least one processor and storing instructions that, when executed by the at least one processor, direct operations, wherein the operations may include: establishing a connection with another communication device; receiving a first message requesting capability information from the other communication device; transmitting a second message including the capability information to the other communication device; receiving configuration information for communication from the other communication device; and performing semantic communication supporting a plurality of tasks based on the configuration information. The configuration information may include information for performing semantic error detection and correction using an attention map generated based on prior information of input data.

[0015] As an example of the present disclosure, a non-transitory computer-readable medium storing at least one instruction includes at least one instruction executable by a processor, wherein the at least one instruction can control a device to establish a connection with another device, receive a first message requesting capability information from the other device, transmit a second message including the capability information to the other device, receive configuration information for communication from the other device, and perform semantic communication supporting a plurality of tasks based on the configuration information. The configuration information can include information for performing semantic error detection and correction using an attention map generated based on prior information of input data.

[0016] The above-described aspects of the present disclosure are only some of the preferred embodiments of the present disclosure, and various embodiments reflecting the technical features of the present disclosure can be derived and understood by a person having ordinary skill in the art based on the detailed description of the present disclosure to be described below.

[0017] The following effects may be achieved by embodiments based on the present disclosure.

[0018] According to the present disclosure, semantic communication can be effectively performed in a wireless communication system.

[0019] The effects that can be obtained from the embodiments of the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly derived and understood by those skilled in the art to which the technical configuration of the present disclosure is applied, from the description of the embodiments of the present disclosure below. In other words, unintended effects that result from implementing the configuration described in the present disclosure can also be derived by those skilled in the art from the embodiments of the present disclosure.

[0020] The accompanying drawings are intended to aid understanding of the present disclosure and, together with detailed descriptions, may provide embodiments of the present disclosure. However, the technical features of the present disclosure are not limited to specific drawings, and the features disclosed in each drawing may be combined with each other to form new embodiments. Reference numerals in each drawing may indicate structural elements.

[0021] Figure 1 illustrates an example of a wireless communication system applicable to the present disclosure.

[0022] FIG. 2 illustrates an example of a wireless device applicable to the present disclosure.

[0023] FIG. 3 illustrates a method for processing a transmission signal applicable to the present disclosure.

[0024] Figure 4 illustrates a communication procedure between a terminal and a base station applicable to the present disclosure.

[0025] FIG. 5 illustrates an example of a communication structure that can be provided in a 6G (6th generation) system applicable to the present disclosure.

[0026] Figure 6 illustrates an electromagnetic spectrum applicable to the present disclosure.

[0027] Figure 7 illustrates a transmitter structure applicable to the present disclosure.

[0028] Figure 8 illustrates an example of a functional framework for application of artificial intelligence technology applicable to the present disclosure.

[0029] Figure 9 illustrates an example of a procedure for utilizing an artificial intelligence model applicable to the present disclosure.

[0030] Figure 10 illustrates a communication procedure based on AI (artificial intelligence) technology applicable to the present disclosure.

[0031] Figure 11 illustrates a communication model applicable to the present disclosure.

[0032] Figure 12 illustrates an example of a semantic communication error.

[0033] Figure 13 shows an example of visualizing the relationships between words in a translation task using the attention technique.

[0034] Figure 14 shows an example of the results of visualizing an attention map for multiple heads for input data in the form of an image.

[0035] FIG. 15 illustrates the functional structure of a source for performing attention map-based semantic error detection and correction considering priors according to one embodiment of the present disclosure.

[0036] FIGS. 16A and 16B illustrate the functional structure of a destination for performing attention map-based semantic error detection and correction considering priors according to one embodiment of the present disclosure.

[0037] FIG. 17 illustrates an example of a procedure for performing semantic communication of a first device according to one embodiment of the present disclosure.

[0038] FIG. 18 illustrates an example of a procedure for performing semantic communication of a second device according to one embodiment of the present disclosure.

[0039] FIG. 19 illustrates an example of a procedure for performing semantic communication using prior information of a first device according to one embodiment of the present disclosure.

[0040] FIG. 20 illustrates an example of a procedure for performing semantic communication using the prior information of a second device according to one embodiment of the present disclosure.

[0041] FIG. 21 illustrates an example of a procedure for supporting attention map-based semantic error detection and correction considering priors according to one embodiment of the present disclosure.

[0042] FIGS. 22a and 22b illustrate examples of a procedure for performing attention map-based semantic error detection and correction considering priors according to one embodiment of the present disclosure.

[0043] FIG. 23 illustrates an example of an initial setup procedure for semantic communication according to one embodiment of the present disclosure.

[0044] Figure 24 shows an example of a final attention map generated using prior information.

[0045] FIG. 25 illustrates an example of a layer structure for generating a prior-based attention map using residual connection and convolution operations according to one embodiment of the present disclosure.

[0046] FIG. 26 illustrates examples of attention maps compared for determining semantic error rates according to one embodiment of the present disclosure.

[0047] Figure 27 illustrates an example of a wireless device applicable to the present disclosure.

[0048] Figure 28 illustrates an example of a portable device applicable to the present disclosure.

[0049] FIG. 29 illustrates an example of a vehicle or autonomous vehicle applicable to the present disclosure.

[0050] Figure 30 illustrates an example of a vehicle applicable to the present disclosure.

[0051] FIG. 31 illustrates an example of an extended reality (XR) device applicable to the present disclosure.

[0052] Figure 32 illustrates an example of a robot applicable to the present disclosure.

[0053] Figure 33 illustrates an example of an AI device applicable to the present disclosure.

[0054] The following embodiments combine components and features of the present disclosure in a predetermined form. Each component or feature may be considered optional unless explicitly stated otherwise. Each component or feature may be implemented without being combined with other components or features. Furthermore, some components and / or features may be combined to form embodiments of the present disclosure. The order of operations described in the embodiments of the present disclosure may be changed. Some components or features of one embodiment may be included in another embodiment or may be replaced with corresponding components or features of another embodiment.

[0055] In the description of the drawings, procedures or steps that may obscure the gist of the present disclosure are not described, and procedures or steps that can be understood by a person skilled in the art are also not described.

[0056] Throughout the specification, when a part is said to "comprising" or "including" a component, this does not mean that other components may be included, but rather that other components may be excluded, unless otherwise specifically stated. In addition, terms such as "...part," "...unit," and "module" described in the specification mean a unit that processes at least one function or operation, which may be implemented by hardware, software, or a combination of hardware and software. In addition, the words "a" or "an," "one," "the," and similar related words may be used in the context of describing the present disclosure (especially in the context of the claims below) to include both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context.

[0057] Embodiments of the present disclosure described herein focus on the data transmission and reception relationship between a base station and a mobile station. Here, the base station is understood as a terminal node of a network that directly communicates with the mobile station. Certain operations described herein as being performed by the base station may, in some cases, be performed by an upper node of the base station.

[0058] That is, in a network consisting of multiple network nodes including a base station, various operations performed for communication with a mobile station may be performed by the base station or other network nodes other than the base station. In this case, the term 'base station' may be replaced by terms such as fixed station, Node B, eNB (eNode B), gNB (gNode B), ng-eNB, advanced base station (ABS), or access point.

[0059] Additionally, in the embodiments of the present disclosure, the term terminal may be replaced with terms such as user equipment (UE), mobile station (MS), subscriber station (SS), mobile subscriber station (MSS), mobile terminal, or advanced mobile station (AMS).

[0060] Additionally, a transmitter refers to a fixed and / or mobile node that provides data or voice services, and a receiver refers to a fixed and / or mobile node that receives data or voice services. Therefore, for uplink, a mobile station can be the transmitter, and a base station can be the receiver. Similarly, for downlink, a mobile station can be the receiver, and a base station can be the transmitter.

[0061] Embodiments of the present disclosure are wireless access systems, such as IEEE 802.xx systems, 3GPP (3 rd Generation Partnership Project) system, 3GPP LTE (Long Term Evolution) system, 3GPP 5G (5 th generation) NR (New Radio) system and 3GPP2 system, and in particular, the embodiments of the present disclosure may be supported by 3GPP TS (technical specification) 38.211, 3GPP TS 38.212, 3GPP TS 38.213, 3GPP TS 38.321 and 3GPP TS 38.331 documents.

[0062] Furthermore, the embodiments of the present disclosure can be applied to other wireless access systems and are not limited to the systems described above. For example, they can also be applied to systems implemented after the 3GPP 5G NR system, and are not limited to a specific system.

[0063] That is, obvious steps or parts not described in the embodiments of the present disclosure can be explained by referring to the above documents. In addition, all terms disclosed in this document can be explained by the above standard documents.

[0064] Hereinafter, preferred embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. The detailed description set forth below, together with the accompanying drawings, is intended to illustrate exemplary embodiments of the present disclosure and is not intended to represent the only embodiments in which the technical configurations of the present disclosure may be implemented.

[0065] Additionally, specific terms used in the embodiments of the present disclosure are provided to aid in understanding of the present disclosure, and the use of such specific terms may be changed to other forms without departing from the technical spirit of the present disclosure.

[0066] The following technology can be applied to various wireless access systems such as CDMA (code division multiple access), FDMA (frequency division multiple access), TDMA (time division multiple access), OFDMA (orthogonal frequency division multiple access), and SC-FDMA (single carrier frequency division multiple access).

[0067]

[0068] For clarity, the following description is based on 3GPP communication systems (e.g., LTE, NR, etc.), but the technical spirit of the present disclosure is not limited thereto. LTE may refer to technology after 3GPP TS 36.xxx Release 8. Specifically, LTE technology after 3GPP TS 36.xxx Release 10 may be referred to as LTE-A, and LTE technology after 3GPP TS 36.xxx Release 13 may be referred to as LTE-A pro. 3GPP NR may refer to technology after TS 38.xxx Release 15. 3GPP 6G may refer to technology after TS Release 17 and / or Release 18. "xxx" refers to a standard document detail number. LTE / NR / 6G may be collectively referred to as a 3GPP system.

[0069] For background information, terms, abbreviations, etc. used in this disclosure, reference may be made to standard documents published prior to this disclosure. For example, reference may be made to the 36.xxx and 38.xxx standard documents.

[0070]

[0071] Wireless communication system applicable to the present disclosure

[0072] Although not limited thereto, the various descriptions, functions, procedures, proposals, methods and / or operational flowcharts of the present disclosure disclosed in this document may be applied to various fields requiring wireless communication / connectivity (e.g., 5G) between devices.

[0073] Hereinafter, more specific examples will be provided with reference to the drawings. In the drawings / descriptions below, the same drawing reference numerals may represent identical or corresponding hardware blocks, software blocks, or functional blocks, unless otherwise described.

[0074] Figure 1 illustrates an example of a wireless communication system applied to the present disclosure.

[0075] Referring to FIG. 1, a wireless communication system (100) applied to the present disclosure includes a wireless device, a base station, and a network. Here, the wireless device refers to a device that performs communication using a wireless access technology (e.g., LTE, LTE-A, LTE-A pro, NR, 5G, 5G-A, 6G) and may be referred to as a communication / wireless / 5G device. Although not limited thereto, the wireless device may include a robot (100a), a vehicle (100b-1, 100b-2), an XR (extended reality) device (100c), a hand-held device (100d), a home appliance (100e), an IoT (Internet of Things) device (100f), and an AI (artificial intelligence) device / server (100g). For example, the vehicle may include a vehicle equipped with a wireless communication function, an autonomous vehicle, a vehicle capable of performing vehicle-to-vehicle communication, etc. Here, the vehicles (100b-1, 100b-2) may include unmanned aerial vehicles (UAVs) (e.g., drones). The XR devices (100c) include augmented reality (AR) / virtual reality (VR) / mixed reality (MR) devices, and may be implemented in the form of head-mounted devices (HMDs), head-up displays (HUDs) installed in vehicles, televisions, smartphones, computers, wearable devices, home appliances, digital signage, vehicles, robots, etc. The portable devices (100d) may include smartphones, smart pads, wearable devices (e.g., smartwatches, smart glasses), computers (e.g., laptops, etc.), etc. The home appliances (100e) may include TVs, refrigerators, washing machines, etc. The IoT devices (100f) may include sensors, smart meters, etc.For example, the base station (120) and the network (130) may also be implemented as wireless devices, and a specific wireless device (120a) may act as a base station / network node to other wireless devices.

[0076] Wireless devices (100a to 100f) can be connected to a network (130) via a base station (120). AI technology can be applied to the wireless devices (100a to 100f), and the wireless devices (100a to 100f) can be connected to an AI server (100g) via a network (130). The network (130) can be configured using a 3G network, a 4G (e.g., LTE) network, a 5G (e.g., NR), or a 6G network. The wireless devices (100a to 100f) can communicate with each other via the base station (120) / network (130), but can also communicate directly (e.g., sidelink communication) without going through the base station (120) / network (130). For example, vehicles (100b-1, 100b-2) can communicate directly (e.g., V2V (vehicle to vehicle) / V2X (vehicle to everything) communication). Additionally, an IoT device (100f) (e.g., a sensor) can communicate directly with another IoT device (e.g., a sensor) or another wireless device (100a to 100f).

[0077] Wireless communication / connection (150a, 150b, 150c) can be established between wireless devices (100a to 100f) / base stations (120), and base stations (120) / base stations (120). Here, the wireless communication / connection can be established through various wireless access technologies such as uplink / downlink communication (150a), sidelink communication (150b) (or D2D communication), and base station-to-base station communication (150c) (e.g., relay, IAB (integrated access backhaul)). Through the wireless communication / connection (150a, 150b, 150c), the wireless device and base station / wireless device, and base stations and base stations can transmit / receive wireless signals to / from each other. For example, the wireless communication / connection (150a, 150b, 150c) can transmit / receive signals through various physical channels. To this end, based on various proposals of the present disclosure, at least some of various configuration information setting processes for transmitting / receiving wireless signals, various signal processing processes (e.g., channel encoding / decoding, modulation / demodulation, resource mapping / demapping, etc.), resource allocation processes, etc. may be performed.

[0078]

[0079] Devices applicable to the present disclosure

[0080] FIG. 2 illustrates an example of a wireless device applicable to the present disclosure.

[0081] Referring to FIG. 2, the wireless device (200) can transmit and receive wireless signals via various wireless access technologies (e.g., LTE, LTE-A, LTE-A pro, NR, 5G, 5G-A, 6G). The wireless device (200) includes at least one processor (202) and at least one memory (204), and may additionally include at least one transceiver (206) and / or at least one antenna (208).

[0082] The processor (202) controls the memory (204) and / or the transceiver (206), and may be configured to implement the descriptions, functions, procedures, proposals, methods, and / or operational flowcharts disclosed in this document. For example, the processor (202) may process information in the memory (204) to generate first information / signal, and then transmit a wireless signal including the first information / signal via the transceiver (206). In addition, the processor (202) may receive a wireless signal including second information / signal via the transceiver (206), and then store information obtained from signal processing of the second information / signal in the memory (204). The memory (204) may be connected to the processor (202) and may store various information related to the operation of the processor (202). For example, the memory (204) may store software code including instructions for performing some or all of the processes controlled by the processor (202), or for performing the descriptions, functions, procedures, proposals, methods, and / or operational flowcharts disclosed herein. Here, the processor (202) and the memory (204) may be part of a communication modem / circuit / chip designed to implement wireless communication technology. The transceiver (206) may be connected to the processor (202) and may transmit and / or receive wireless signals via at least one antenna (208). The transceiver (206) may include a transmitter and / or a receiver. The transceiver (206) may be used interchangeably with an RF (radio frequency) unit. In the present disclosure, a wireless device may also mean a communication modem / circuit / chip.

[0083] Hereinafter, the hardware elements of the wireless device (200) will be described in more detail. Although not limited thereto, at least one protocol layer may be implemented by at least one processor (202). For example, at least one processor (202) may implement at least one layer (e.g., a functional layer such as physical (PHY), media access control (MAC), radio link control (RLC), packet data convergence protocol (PDCP), radio resource control (RRC), and service data adaptation protocol (SDAP)). At least one processor (202) may generate at least one Protocol Data Unit (PDU) and / or at least one Service Data Unit (SDU) according to the descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document. At least one processor (202) may generate a message, control information, data, or information according to the descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document. At least one processor (202) can generate a signal (e.g., a baseband signal) including a PDU, an SDU, a message, control information, data or information according to the functions, procedures, proposals and / or methods disclosed in this document, and provide the signal to at least one transceiver (206). At least one processor (202) can receive a signal (e.g., a baseband signal) from at least one transceiver (206) and obtain the PDU, SDU, message, control information, data or information according to the descriptions, functions, procedures, proposals, methods and / or operational flowcharts disclosed in this document.

[0084] At least one processor (202) may be referred to as a controller, a microcontroller, a microprocessor, or a microcomputer. The at least one processor (202) may be implemented by hardware, firmware, software, or a combination thereof. For example, at least one application specific integrated circuit (ASIC), at least one digital signal processor (DSP), at least one digital signal processing device (DSPD), at least one programmable logic device (PLD), or at least one field programmable gate array (FPGA) may be included in the at least one processor (202). The descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document may be implemented using firmware or software, and the firmware or software may be implemented to include modules, procedures, functions, etc. The descriptions, functions, procedures, proposals, methods, and / or operation flowcharts disclosed in this document may be included in the at least one processor (202), or may be stored in at least one memory (204) and executed by the at least one processor (202). The descriptions, functions, procedures, suggestions, methods and / or flowcharts disclosed in this document may be implemented using firmware or software in the form of code, instructions and / or sets of instructions.

[0085] At least one memory (204) can be connected to at least one processor (202) and can store various forms of data, signals, messages, information, programs, codes, instructions and / or commands. The at least one memory (204) can be configured as a read only memory (ROM), a random access memory (RAM), an erasable programmable read only memory (EPROM), a flash memory, a hard drive, a register, a cache memory, a computer readable storage medium and / or a combination thereof. The at least one memory (204) can be located internally and / or externally to the at least one processor (202). In addition, the at least one memory (204) can be connected to the at least one processor (202) via various technologies such as a wired or wireless connection.

[0086] At least one transceiver (206) can transmit user data, control information, wireless signals / channels, etc., mentioned in the methods and / or flowcharts of this document to at least one other device. At least one transceiver (206) can receive user data, control information, wireless signals / channels, etc. mentioned in the descriptions, functions, procedures, proposals, methods and / or flowcharts disclosed in this document from at least one other device. For example, at least one transceiver (206) can be connected to at least one processor (202) and can transmit and receive wireless signals. For example, at least one processor (202) can control at least one transceiver (206) to transmit user data, control information, or wireless signals to at least one other device. Furthermore, at least one processor (202) can control at least one transceiver (206) to receive user data, control information, or wireless signals from at least one other device. In addition, at least one transceiver (206) may be connected to at least one antenna (208), and at least one transceiver (206) may be configured to transmit and receive user data, control information, wireless signals / channels, etc. mentioned in the descriptions, functions, procedures, proposals, methods and / or operation flowcharts disclosed in this document through at least one antenna (208). In this document, at least one antenna may be a plurality of physical antennas or a plurality of logical antennas (e.g., antenna ports). At least one transceiver (206) may convert the received wireless signals / channels, etc. from RF band signals to baseband signals in order to process the received user data, control information, wireless signals / channels, etc. using at least one processor (202). At least one transceiver (206) may convert the processed user data, control information, wireless signals / channels, etc. from baseband signals to RF band signals using at least one processor (202).For this purpose, at least one transceiver (206) may include an (analog) oscillator and / or filter.

[0087] The components of the wireless device described with reference to FIG. 2 may be referred to by different terms in terms of functionality. For example, the processor (202) may be referred to as a control unit, the transceiver (206) as a communication unit, and the memory (204) as a storage unit. In some cases, the communication unit may be used to mean at least a portion of the processor (202) and the transceiver (206).

[0088] The structure of the wireless device described with reference to FIG. 2 can be understood as the structure of at least a portion of various devices. For example, the structure of the wireless device illustrated in FIG. 2 can be at least a portion of various devices described with reference to FIG. 1 (e.g., a robot (100a), a vehicle (100b-1, 100b-2), an XR device (100c), a portable device (100d), a home appliance (100e), an IoT device (100f), an AI device / server (100g)). Furthermore, according to various embodiments, in addition to the components illustrated in FIG. 2, the device may further include other components.

[0089] For example, the device may be a portable device such as a smartphone, a smart pad, a wearable device (e.g., a smart watch, smart glasses), or a portable computer (e.g., a laptop, etc.). In this case, the device may further include at least one of a power supply unit that supplies power and includes a wired / wireless charging circuit, a battery, etc., an interface unit that includes at least one port for connection with another device (e.g., an audio input / output port, a video input / output port), and an input / output unit for inputting and outputting image information / signals, audio information / signals, data, and / or information input from a user.

[0090] For example, the device may be a mobile device such as a mobile robot, a vehicle, a train, an aerial vehicle (AV), a ship, etc. In this case, the device may further include at least one of a driving unit including at least one of an engine, a motor, a power train, wheels, brakes, and a steering unit of the device, a power supply unit including a wired / wireless charging circuit, a battery, etc. that supplies power, a sensor unit that senses status information, environmental information, and user information of the device or its surroundings, an autonomous driving unit that performs functions such as path maintenance, speed control, and destination setting, and a position measurement unit that obtains location information of the mobile device through a global positioning system (GPS) and various sensors.

[0091] For example, the device may be an XR device such as an HMD, a head-up display (HUD) installed in a vehicle, a television, a smartphone, a computer, a wearable device, a home appliance, a digital signage, a vehicle, a robot, etc. In this case, the device may further include at least one of a power supply unit that supplies power and includes a wired / wireless charging circuit, a battery, etc., an input / output unit that obtains control information, data, etc. from the outside and outputs the generated XR object, and a sensor unit that senses status information, environmental information, and user information of the device or the surroundings of the device.

[0092] For example, the device may be a robot that can be classified into industrial, medical, household, military, etc. types depending on the purpose or field of use. In this case, the device may further include at least one of a sensor unit that senses status information, environmental information, and user information of the device or its surroundings, and a driving unit that performs various physical actions, such as moving the robot joints.

[0093] For example, the device may be an AI device such as a TV, a projector, a smartphone, a PC, a laptop, a digital broadcasting terminal, a tablet PC, a wearable device, a set-top box (STB), a radio, a washing machine, a refrigerator, digital signage, a robot, a vehicle, etc. In this case, the device may further include at least one of an input unit that acquires various types of data from the outside, an output unit that generates output related to sight, hearing, or touch, a sensor unit that senses status information, environmental information, and user information of the device or its surroundings, and a training unit that trains a model composed of an artificial neural network using learning data.

[0094] The structure of the wireless device illustrated in FIG. 2 may be understood as a part of a RAN node (e.g., a base station, DU, RU, RRH, etc.). That is, the device illustrated in FIG. 2 may be a RAN node. In this case, the device may further include a wired transceiver for front haul and / or back haul communications. However, if the front haul and / or back haul communications are based on wireless communications, at least one transceiver (206) illustrated in FIG. 2 may be used for front haul and / or back haul communications, and a wired transceiver may not be included.

[0095]

[0096] FIG. 3 illustrates a method for processing a transmission signal applicable to the present disclosure. For example, the transmission signal may be processed by a signal processing circuit. At this time, the signal processing circuit (300) may include scramblers (310), modulators (320), a layer mapper (330), a precoder (340), resource mappers (350), and signal generators (360). At this time, for example, the operation / function of FIG. 3 may be performed in the processor (202) and / or the transceiver (206) of FIG. 2. Furthermore, for example, the hardware elements of FIG. 3 may be implemented in the processor (202) and / or the transceiver (206) of FIG. 2. For example, blocks 310 to 360 may be implemented in the processor (202) of FIG. 2. Additionally, blocks 310 to 350 may be implemented in the processor (202) of FIG. 2, and block 360 may be implemented in the transceiver (206) of FIG. 2, and are not limited to the above-described embodiment.

[0097] The codeword can be converted into a wireless signal through the signal processing circuit (300) of FIG. 3. Here, the codeword is an encoded bit sequence of an information block. The information block may include a transport block (e.g., a UL-SCH transport block, a DL-SCH transport block). Here, the information block may include data related to AI (e.g., training data, AI model data, input data, output data, etc.), and the codeword may be an encoded bit sequence corresponding to the data related to AI. The wireless signal may be transmitted through various physical channels (e.g., a PUSCH, a PDSCH). Specifically, the codeword may be converted into a bit sequence scrambled by scramblers (310). The scramble sequence used for scrambling is generated based on an initialization value, and the initialization value may include ID information of the wireless device, etc. The scrambled bit sequence may be modulated into a modulation symbol sequence by modulators (320). Modulation schemes may include pi / 2-BPSK (pi / 2-binary phase shift keying), m-PSK (m-phase shift keying), m-QAM (m-quadrature amplitude modulation), etc.

[0098] A complex modulation symbol sequence can be mapped to at least one transport layer by a layer mapper (330). Here, a transport layer is a logical resource unit for mapping a signal or data transmitted through spatial resources to antenna ports, and one transport layer can correspond to one stream or one antenna port. Each of the complex modulation symbols included in the complex modulation symbol sequence is mapped to at least one transport layer, thereby determining which antenna port it will be transmitted through. The modulation symbols of each transport layer can be mapped to the corresponding antenna port(s) by a precoder (340). The output z of the precoder (340) can be obtained by multiplying the output y of the layer mapper (330) by a precoding matrix W of NХM. Here, N is the number of antenna ports, and M is the number of transport layers. Here, the precoder (340) may perform precoding after performing transform precoding (e.g., discrete Fourier transform (DFT) transform) on complex modulation symbols. Additionally, the precoder (340) may perform precoding without performing transform precoding.

[0099] Resource mappers (350) can map modulation symbols of each antenna port to time-frequency resources. The time-frequency resources may include a plurality of symbols (e.g., CP-OFDMA symbols, DFT-s-OFDMA symbols) in the time domain and a plurality of subcarriers in the frequency domain. Signal generators (360) generate wireless signals from the mapped modulation symbols, and the generated wireless signals can be transmitted to other devices through each antenna. To this end, each of the signal generators (360) may include an inverse fast Fourier transform (IFFT) module, a cyclic prefix (CP) inserter, a digital-to-analog converter (DAC), a frequency uplink converter, etc.

[0100] The signal processing process for a received signal in a wireless device may be configured in reverse order of the signal processing process (310 to 360) of FIG. 3. For example, a wireless device (e.g., 200 of FIG. 2) may receive a wireless signal from the outside through an antenna port / transceiver. The received wireless signal may be converted into a baseband signal through a signal restorer. For this purpose, the signal restorer may include a frequency downlink converter, an analog-to-digital converter (ADC), a CP remover, and a fast Fourier transform (FFT) module. Thereafter, the baseband signal may be restored to a codeword through a resource demapper process, a postcoding process, a demodulation process, and a descrambling process. The codeword may be restored to the original information block through decoding. Therefore, a signal processing circuit (not shown) for a received signal may include a signal restorer, a resource demapper, a postcoder, a demodulator, a descrambler, and a decoder.

[0101] The signal processing circuit (300) described with reference to FIG. 3 is exemplified as including a plurality of scramblers (310), modulators (320), a plurality of resource mappers (350), and a plurality of signal generators (360). However, at least one of the scramblers, modulators, resource mappers, and signal generators may be implemented as a single integrated structure. That is, the number of at least one of the scramblers, modulators, resource mappers, and signal generators may be smaller than the number of layers. Furthermore, at least one of the components exemplified in FIG. 3 may be omitted.

[0102]

[0103] Figure 4 illustrates a communication procedure between a terminal and a base station applicable to the present disclosure. Figure 4 illustrates operations of a terminal (410) and a base station (420) transmitting and / or receiving data and operations performed prior thereto.

[0104] Referring to FIG. 4, in step 401, the terminal (410) and the base station (420) perform synchronization. For example, the terminal (410) performs an initial cell search operation. Specifically, the terminal (410) can detect at least one synchronization signal transmitted from the base station (420) according to a predefined rule. Here, the synchronization signal can include multiple synchronization signals classified according to structure or purpose (e.g., primary synchronization signal, secondary synchronization signal). Through this, the terminal (410) can check the boundary of the frame, subframe, slot, and / or symbol of the base station (420) and obtain information about the base station (420) (e.g., cell identifier).

[0105] In step 403, the terminal (410) obtains system information transmitted from the base station (420). The system information is information related to the properties, characteristics, and / or capabilities of the base station (420) required to access the base station (420) and use the service, and may be classified by content (e.g., whether it is essential for access), transmission structure (e.g., channel used, whether provided on-demand), etc., and may be classified into, for example, a master information block (MIB) and a system information block (SIB). If necessary, the terminal (410) may transmit a signal requesting system information before receiving the system information. The system information may include information related to an AI function. For example, the system information may include at least one of information related to an AI model, information related to training, and information related to inference / prediction, as information required for operations performed based on AI. However, the request and provision of the system information may be performed after a random access procedure described below.

[0106] In step 405, the terminal (410) and the base station (420) perform a random access procedure. The terminal (410) may transmit and / or receive at least one message (e.g., a random access preamble, a random access response (RAR) message, etc.) for the random access procedure based on information related to the random access channel of the base station (420) obtained through system information (e.g., channel position, channel structure, supported preamble structure, etc.). For example, the terminal (410) may transmit a preamble (e.g., MSG1) through the random access channel, receive an RAR message (e.g., MSG2), transmit a message (e.g., MSG3) including information related to the terminal (410) (e.g., identification information) to the base station (420) using scheduling information included in the RAR message, and receive a message (e.g., MSG4) for contention resolution and / or connection establishment. As another example, MSG1 and MSG3 may be sent and received as one message, or MSG2 and MSG4 may be sent and received as one message.

[0107] In step 407, the terminal (410) and the base station (420) perform signaling of control information. Here, the control information may be defined in various layers, such as a layer that controls a connection (e.g., a radio resource control (RRC) layer), a layer that handles mapping between logical channels and transport channels (e.g., a media access control (MAC) layer), and a layer that handles physical channels (e.g., a physical (PHY) layer). For example, the terminal (410) and the base station (420) may perform at least one of signaling for establishing a connection, signaling for determining settings related to communication, and signaling for indicating allocated resources. In addition, the signaling of the control information may be performed to convey information related to an AI function. For example, the information related to an AI function is information necessary for an operation performed based on AI, and may include at least one of information related to an AI model, information related to training, and information related to inference / prediction. More specifically, information related to the AI ​​function signaled in step 407 may be combined and / or linked with information related to the AI ​​function signaled in step 403, and the two may be defined in a hierarchical, mutually complementary, or substitutive structure.

[0108] In step 409, the terminal (410) and the base station (420) transmit and / or receive data. In other words, the terminal (410) and the base station (420) can process, transmit, and / or receive data based on the signaling of the control information. For example, when transmitting data, the terminal (410) or the base station (420) can perform at least one of channel encoding, rate matching, scrambling, constellation mapping, layer mapping, waveform modulation, antenna mapping, and resource mapping on the information bits. Conversely, when receiving data, the terminal (410) or the base station (420) can perform at least one of signal extraction from resources, waveform demodulation for each antenna, signal arrangement considering layer mapping, constellation demapping, descrambling, and channel decoding. Here, the transmitted data is data related to AI, and may include, for example, data for AI-based operations or data generated by AI-based operations.

[0109] Steps 401 to 409 illustrated with reference to FIG. 4 do not necessarily have to be performed in the order illustrated in FIG. 4, and the order of at least some of the steps may vary. Furthermore, at least some of steps 401 to 409 may be combined into a single step or omitted. That is, the steps illustrated in FIG. 4 may be performed in various modified forms.

[0110]

[0111] 6G wireless communication systems and core implementation technologies of 6G systems

[0112] The 5G system defines various operating bands within FR1 (frequency range 1), which covers 410 MHz to 7125 MHz, and FR2 (frequency range 2), which covers 24,250 MHz to 71,000 MHz. Various frequencies are being discussed as operating bands for the subsequent 6G system, and the use of higher frequencies than 5G systems is also being considered for wider bandwidth and higher transmission speeds. One such band is the THz (terahertz) frequency band, which covers approximately 100 GHz to 10 THz. The THz frequency band is a band that has both the transparency of radio waves and the straightness of light waves, and communications using the THz frequency band are expected to play a transitional role from existing radio-centered communications to lightwave-based communications.

[0113] 6G systems utilizing the THz frequency band have the following goals: i) very high data rates per device, ii) a very large number of connected devices, iii) global connectivity, iv) very low latency, v) reduced energy consumption of battery-free IoT devices, vi) ultra-reliable connectivity, and vii) connected intelligence with machine learning capabilities. The vision of the 6G system can be divided into four aspects: “intelligent connectivity,” “deep connectivity,” “holographic connectivity,” and “ubiquitous connectivity,” and the 6G system can be designed to satisfy the requirements as shown in [Table 1] below.

[0114] Per device peak data rate1 TbpsE2E latency1 msMaximum spectral efficiency100 bps / HzMobility supportup to 1000 km / hrSatellite integrationFullyAIFullyAutonomous vehicleFullyXRFullyHaptic CommunicationFully

[0115] At this time, the 6G system may have key factors such as enhanced mobile broadband (eMBB), ultra-reliable low latency communications (URLLC), massive machine type communications (mMTC), AI integrated communication, tactile internet, high throughput, high network capacity, high energy efficiency, low backhaul and access network congestion, and enhanced data security. FIG. 5 illustrates an example of a communication structure that can be provided in a 6G system applicable to the present disclosure. Referring to FIG. 5, the 6G system is expected to have 50 times higher simultaneous wireless communication connectivity than a 5G wireless communication system. URLLC, a key feature of 5G, is expected to become an even more crucial technology in 6G communications, offering end-to-end latency of less than 1 ms. Furthermore, 6G systems will boast significantly higher volumetric spectral efficiency than the commonly used area spectral efficiency. 6G systems can offer extremely long battery life and advanced battery technologies for energy harvesting, eliminating the need for separate charging for mobile devices in 6G systems. New network characteristics in 6G may include:

[0116] - Satellite integrated network: 6G is expected to integrate with satellites to provide a global mobile network. The integration of terrestrial, satellite, and airborne networks into a single wireless communications system is crucial for 6G.

[0117] Connected Intelligence: Unlike previous generations of wireless communication systems, 6G is revolutionary, upgrading the wireless evolution from "connected objects" to "connected intelligence." AI can be applied at every stage of the communication process (or at every signal processing step, as described below).

[0118] - Seamless integration of wireless information and energy transfer: 6G wireless networks will transfer power to charge the batteries of devices such as smartphones and sensors. Therefore, wireless information and energy transfer (WIET) will be integrated.

[0119] - Ubiquitous super 3D connectivity: Access to networks and core network functions of drones and very low Earth orbit satellites will create super 3D connectivity in 6G ubiquitous.

[0120] Some general requirements for the new network characteristics of 6G, such as the above, may be as follows:

[0121] - Small cell networks: The concept of small cell networks was introduced to improve received signal quality in cellular systems by increasing throughput, energy efficiency, and spectral efficiency. Consequently, small cell networks are essential for 5G and beyond 5G (5GB) wireless communication systems. Accordingly, 6G wireless communication systems also adopt the characteristics of small cell networks.

[0122] Ultra-dense heterogeneous networks: Ultra-dense heterogeneous networks will be another key feature of 6G wireless communication systems. Multi-tier networks comprised of heterogeneous networks improve overall QoS and reduce costs.

[0123] High-capacity backhaul: Backhaul connections are characterized by high-capacity backhaul networks to support high-volume traffic. High-speed fiber optics and free-space optics (FSO) systems may be potential solutions to this problem.

[0124] - Radar technology integrated with mobile technology: High-precision localization (or location-based services) through communications is a key feature of 6G wireless communication systems. Therefore, radar systems will be integrated with 6G networks.

[0125] - Softwarization and virtualization: Softwarization and virtualization are two critical features that form the foundation of the design process for 5GB networks to ensure flexibility, reconfigurability, and programmability. Furthermore, billions of devices can be shared on a shared physical infrastructure.

[0126] To satisfy the above-mentioned characteristics, the core implementation technologies of the 6G system may include artificial intelligence (AI), THz (terahertz) communication, optical wireless technology, FSO backhaul network, massive MIMO technology, blockchain, 3D networking, quantum communication, unmanned aerial vehicles, cell-free communication, wireless information and energy transfer (WIET), integration of sensing and communication, integration of access backhaul networks, holographic beamforming, big data analysis, and large intelligent surface (LIS).

[0127] For example, THz communication can be utilized in 6G systems. THz communication is a communication that utilizes a spectrum in a frequency band between 0.3 THz and 3 THz with a corresponding wavelength in the range of 0.1 mm to 1 mm, as shown in FIG. 6. Referring to FIG. 6, the frequency band of THz waves is located in the middle region between the infrared band and the millimeter wave band, and therefore, THz waves can be understood as radio waves with the shortest wavelength and light waves with the longest wavelength. Therefore, THz waves share some of the characteristics of infrared and microwave waves, and specifically, they can simultaneously have the transparency of electromagnetic waves and the straightness of light waves.

[0128]

[0129] Figure 7 illustrates a transmitter structure applicable to the present disclosure.

[0130] Referring to Figure 7, in order to modulate data into an optical signal, an optical source of a laser can be passed through an optical wave guide to change the phase of the signal, etc. At this time, data is loaded by changing the electrical characteristics through a microwave contact, etc. Therefore, the optical modulator output is formed as a modulated waveform.

[0131] Data may be provided from a data signal generator. Here, the data may include various user data, configuration information, control information, etc. transmitted through a channel. Furthermore, the data may include data related to AI-based operations, such as information for configuring an AI model, input / output data for tasks of the AI ​​model, etc. To this end, components related to AI functions (e.g., an AI processing unit) may be included in the data signal generator or may be linked to the data signal generator.

[0132] An optical / electronic converter (O / E converter) can generate THz pulses by optical rectification using a nonlinear crystal, photoelectric conversion using a photoconductive antenna, or emission from a bunch of relativistic electrons. The THz pulse generated in the above manner can have a length in the range of femtoseconds to picoseconds. The optical / electronic converter (O / E converter) performs down conversion by utilizing the nonlinearity of the device.

[0133] Considering the THz spectrum usage, it is likely that multiple contiguous GHz bands will be used for THz systems, either fixed or for mobile services. For an outdoor scenario, the available bandwidth can be categorized based on an oxygen attenuation of 10^2 dB / km in the spectrum up to 1 THz. Accordingly, a framework in which the available bandwidth is divided into multiple band chunks can be considered. As an example of this framework, if the THz pulse length for a single carrier is set to 50 ps, ​​the bandwidth (BW) becomes approximately 20 GHz.

[0134] Effective down-conversion from the infrared band to the THz band depends on how to utilize the nonlinearity of the optical / electrical converter (O / E converter). In other words, to down-convert to the desired THz band, it is necessary to design an O / E converter with the most ideal non-linearity for transferring to the THz band. If an O / E converter that is not suitable for the target frequency band is used, errors in the amplitude and phase of the pulse are likely to occur.

[0135] A THz transmission and reception system can be implemented using a single optical-to-electrical converter in a single-carrier system. Depending on the channel environment, optical-to-electrical converters may be required as many as the number of carriers in a multi-carrier system. This phenomenon will be particularly noticeable in a multi-carrier system that utilizes multiple broadbands according to the aforementioned spectrum usage plan. In this regard, a frame structure for the multi-carrier system may be considered. A signal down-frequency converted based on an optical-to-electrical converter may be transmitted in a specific resource region (e.g., a specific frame). The frequency region of the specific resource region may include multiple chunks. Each chunk may be composed of at least one component carrier (CC).

[0136]

[0137] 6G systems may introduce AI technology. Efficient resource management and optimization are necessary to maintain connectivity between various services and devices. AI technology may include technologies that perform data analysis, pattern recognition, and predictive modeling using AI / ML (artificial intelligence / machine learning) models. Here, an AI / ML model can be understood as a set of parameter values ​​and / or weight values ​​related to mathematical formulas or algorithms generated through learning to discover patterns in input data or make predictions. To create such an AI / ML model, an AI / ML model learning process is required, which builds an AI / ML model by learning the relationship between inputs and outputs in a data-driven manner. Various learning algorithms, such as supervised learning, unsupervised learning, and reinforcement learning, can be utilized. To generate output, a user can input specific data into a trained AI / ML model, and the process of obtaining output data by inputting input data into an AI / ML model can be referred to as AI / ML "inference" or "prediction."

[0138] Network control parameters can be obtained as output through AI / ML inference using trained AI / ML models. Users can utilize the output parameter values ​​to improve network efficiency. For example, AI technology can be utilized in various fields, such as wireless network resource allocation, traffic management, fault prediction, and quality of service (QoS) management. In particular, machine learning can efficiently allocate resources even in dynamically changing network environments based on real-time data. Therefore, AI technology can be utilized to provide hyper-connectivity and ultra-low latency.

[0139] At this time, AI / ML inference can be performed based on a combination of various devices. For example, the UE and the network can jointly perform AI / ML inference, and such an AI / ML model can be referred to as a two-sided AI / ML model or a two-sided model. In this case, the UE can perform the first part of the inference first, and the base station can perform the remaining inference, or vice versa. Alternatively, inference can be performed entirely on the UE, and such an AI / ML model can be referred to as a UE-side AI / ML model or a UE-side model.

[0140] Additionally, life cycle management (LCM) can be performed for AI / ML models. Life cycle management can include model training, model deployment, model inference, model monitoring, and model updates. This may require support for data collection, model training, functional / model identification, model delivery / transfer, model inference operations, functional / model selection / activation / deactivation / fallback, functional / model monitoring, model updates, and UE capabilities.

[0141] For example, AI / ML technology can be operated based on a functional framework such as FIG. 8. FIG. 8 illustrates an example of a functional framework for application of AI / ML technology applicable to the present disclosure. First, a data collection function (810) performs data preparation on input data collected from objects (e.g., UE, RAN node, network node, etc.) to generate training data (801), monitoring data (803), and / or inference data (805) including processed input data. A model training function (820), which receives training data (801) from the data collection function (810), performs training on an AI / ML model using the training data (801) and provides a trained / updated model (813) to a model repository (840). The model repository (840) can store and retain the received trained / updated model (813).

[0142] A management function (830) may be utilized to control AI / ML model training. The management function (830) may control the operation of the AI / ML model or AI / ML functions, or supervise their performance. To this end, the management function (830) may receive monitoring data (830) from the data collection function (810) and inference output (809) from the inference function (840). The management function (830) manages the data received from the data collection function (810) and the inference function (840) so that the inference task can be performed efficiently. That is, the management function (830) may transmit performance feedback or a retraining request (807) to the model training function (820) to improve the inference task. Here, the performance feedback may be utilized to indicate a learning goal or as a reward for reinforcement learning. Additionally, the management function (830) can transmit management instructions (811) that instruct the inference function (840) to select AI / ML models or AL / ML-based functions to be used, activate / deactivate them, or switch to non-AI / ML operation.

[0143] The inference function (840) generates an inference output (809) by performing inference and / or prediction using the inference data (805) received by the data collection function (810). Here, the inference output (809) refers to the inference output of the AI / ML model used by the inference function (840), and the details of the inference output may vary depending on the use case. The AI / ML model used by the inference function (840) can be controlled by the management function (830). That is, the management function (830) can transmit a model transfer / forward request signal (815) to request a necessary AI / ML model to the model repository function (850), and the model repository function (850) can transmit the corresponding AI / ML model to the inference function (840) via a model transfer / forward signal (817). Therefore, the inference function (840) can perform inference using the AI / ML model (817) according to the received management instruction (811).

[0144] Additionally, the management function (830) may trigger or perform a designated task / action based on the inference output (809). Accordingly, the management function (830) may trigger a task / action for another entity (e.g., at least one UE, at least one RAN node, at least one network node, etc.) or for itself. Any one of the functions exemplified in FIG. 8 described above may be performed by two or more entities, including the RAN, the network node, the network operator's OAM, or the UE, in collaboration. This may be referred to as a split AI operation.

[0145] Not all of the functions (810 to 850) illustrated in FIG. 8 need to be used to utilize the AI / ML model, and the method of combining them is not limited to a specific method. Accordingly, the functions (810 to 850) may be operated in an integrated manner, or some functions may be omitted. Furthermore, the functions (810 to 850) illustrated in FIG. 8 are not necessarily limited to being implemented as separate devices or apparatuses. For example, some or all of the functions (810 to 850) may be included in the processor (202) of FIG. 2. Furthermore, the model storage function (850) may be included in the memory (204) of FIG. 2.

[0146]

[0147] FIG. 9 illustrates an example of a procedure for utilizing an AI model applicable to the present disclosure. FIG. 9 illustrates a case where a model training function (820) is included in a network node and a model inference function (840) is included in a RAN node. Referring to FIG. 9, in step 1, RAN node 1 and RAN node 2 transmit input data (e.g., training data) for training an AI model to the network node. Here, RAN node 1 and RAN node 2 may transmit data collected from the UE (e.g., measurements of the UE related to RSRP, RSRQ, SINR of the serving cell and neighboring cells, the UE's position, speed, etc.) together to the network node. In step 2, the network node trains the AI ​​model using the received training data. In step 3, the network node distributes / updates the AI ​​model to RAN node 1 and / or RAN node 2. RAN node 1 and / or RAN node 2 may also continue model training based on the received AI model. In this procedure, it is assumed that the AI ​​model is deployed / updated only to RAN node 1. In step 4, RAN node 1 receives input data (e.g., inference data) for AI model inference from UE and RAN node 2. In step 5, RAN node 1 performs AI model-based inference using the received inference data to generate output data (e.g., prediction or decision). In step 6, if applicable, RAN node 1 may transmit model performance feedback to network nodes. In step 7, RAN node 1, RAN node 2, and UE (or 'RAN node 1 and UE', or 'RAN node 1 and RAN node 2') perform actions based on the output data. For example, in case of load balancing operation, the UE may move from RAN node 1 to RAN node 2. In step 8, RAN node 1 and RAN node 2 transmit feedback information to network nodes.

[0148] Network nodes can manage AI models based on feedback information regarding the AI ​​model's inference results. For example, the network node can perform additional training on the AI ​​model or generate additional information about the AI ​​model (e.g., performance information, accuracy information, etc.). If additional training is performed on the AI ​​model, the network node can distribute the updated AI model to RAN node 1.

[0149] As described with reference to Figure 9, model training can be performed by network nodes, and inference using the model can be performed by RAN node 1. In other words, the model training and inference functions can be distributed. Typically, model training requires significant computational resources because it utilizes large amounts of data and complex algorithms for optimization. In contrast, inference, which uses a trained model to derive conclusions about new data, requires relatively fewer computational resources compared to model training. Therefore, using the procedure of Figure 9, model training can be performed through network nodes when the computational resources of the UE or RAN node are insufficient. Furthermore, security can be ensured for the AI ​​model because the AI ​​model is not disclosed to the UE.

[0150] FIG. 9 illustrates a case where a model training function (820) is included in a network node and a model inference function (840) is included in a RAN node, but the present disclosure is not limited thereto. For example, if the computational resources of the RAN node are sufficient, both the model training function (820) and the model inference function (840) may be included in RAN node 1. In this case, RAN node 1 receives training data for training an AI model from the UE and RAN node 2. RAN node 1 trains the AI ​​model using the received training data. Thereafter, RAN node 1 receives inference data for AI model inference from the UE and RAN node 2. RAN node 1 performs inference based on the AI ​​model using the received inference data to generate output data. Based on the output data, the UE, RAN node 1, and RAN node 2 may perform operations related to communication (e.g., handover, cell change). Thereafter, the UE and RAN node 2 may transmit feedback regarding the operations to RAN node 1. Therefore, RAN node 1 can learn the AI ​​model and update its own AI model using feedback information regarding the AI ​​model's inference results. According to the aforementioned method, signaling with the network is not required for AI model learning and inference, thereby reducing network load and delays until the AI ​​model is trained or inference results are received. Furthermore, since the UE performs inference using the AI ​​model, its personal information is not transmitted to network nodes, etc., thereby enhancing the security of personal information.

[0151] As another example, the model training function (820) may be included in the RAN node, and the model inference function (840) may be included in the UE. The RAN node receives training data for training an AI model from the UE, and trains the AI ​​model using the received training data. The RAN node distributes the trained AI model to the UE. The UE may perform inference based on the received AI model to generate output data. At this time, the data for inference may be received from the RAN node, or the UE may use data acquired on its own. The UE and the RAN node may perform communication-related operations based on the output data generated by inference. Thereafter, the UE may transmit feedback regarding the operation to the RAN node. Therefore, the RAN node may train the AI ​​model and distribute the updated AI model to the UE through feedback information regarding the inference result of the AI ​​model. According to the above-described method, the model training function (820) may be included in the RAN node, and the model inference function (840) may be included in the UE to perform inference, thereby reducing the load on the RAN node. Additionally, if the UE uses data acquired on its own to make inferences, even if the UE loses connection with the RAN node after receiving the AI ​​model, the UE can continue to infer the distributed AI model and perform operations based on the inference results.

[0152] According to the framework and procedures described above, an AI model can be trained and utilized in a wireless communication system. The model training function (820) and the model inference function (840) can be combined in various ways and are not necessarily limited to the case of FIG. 9. In the framework and procedures described above, various types of data, such as input data, training data, and inference data, are introduced, and the specific content of the data described above may vary depending on the task for which the AI ​​model is utilized. For example, information used in various embodiments of the present disclosure described below may be included in the data described above.

[0153]

[0154] Figure 10 illustrates an AI technology-based communication procedure applicable to the present disclosure. The detailed procedures illustrated in Figure 10 can be combined with various embodiments of the present disclosure described below. For example, data generated according to various embodiments of the present disclosure can be used for operations (e.g., configuration, training, inference, and / or data transmission / reception) in at least one of the detailed procedures illustrated in Figure 10. As another example, the results of the inference illustrated in Figure 10 can be used to transmit and / or receive data according to various embodiments of the present disclosure.

[0155] Referring to FIG. 10, in step S1001, at least one of the UE (1010), the RAN node (1020), and the network node (1030) performs an initial access procedure. For example, in this step, at least one of an initial cell search operation, a system information acquisition operation, a random access operation, and a registration operation may be performed. In step S1003, at least one of the UE (1010), the RAN node (1020), and the network node (1030) performs a configuration procedure. Through the configuration procedure, parameters, resources, connections, and / or entities necessary for performing subsequent procedures in layers between the UE (1010) and the RAN node (1020) and / or in at least one layer between the UE (1010) and the network node (1030) may be determined and / or created. In this case, the configuration procedure may be performed based on information, status, and / or characteristics of an AI model used for subsequent training and inference.

[0156] In step S1005, at least one of the UE (1010), the RAN node (1020), and the network node (1030) performs a model training procedure. At least one of the UE (1010), the RAN node (1020), and the network node (1030) may collect training data and perform learning using the training data. For example, the model training procedure may be performed as described with reference to FIG. 9. If an offline-trained model is used, this step may be omitted.

[0157] In step S1007, at least one of the UE (1010), the RAN node (1020), and the network node (1030) performs a task using the trained model. That is, the task may be performed based on the results of inference and / or prediction using the trained model. For example, the task may be a procedure belonging to a communication protocol, and may be a preparatory operation for subsequent data transmission and / or reception, or may be related to data transmission and / or reception, or may be related to data processing (e.g., encoding, decoding, etc.).

[0158] In step S1009, at least one of the UE (1010), the RAN node (1020), and the network node (1030) transmits and / or receives data. At this time, the result of the task performed in step 1007 may be used. In some cases, the task performed in step 1007 may include transmitting and / or receiving data, in which case this step may be omitted as it is part of step 1007.

[0159]

[0160] Specific embodiments of the present disclosure

[0161] The present disclosure relates to semantic communication in a wireless communication system, and more particularly, to a technique for detecting and correcting errors occurring in a semantic representation (SR) transmitted and received for semantic communication. More specifically, the present disclosure proposes various embodiments for detecting and correcting semantic errors using an attention map generated based on input data and prior information related to the data in a system supporting semantic communication.

[0162]

[0163] Problems related to communication can be divided into three levels, as illustrated in Figure 11, according to the philosophy of Shannon and Weaver. Figure 11 illustrates a communication model applicable to the present disclosure. The problem at Level A (1110) is a technical problem related to how accurately symbols in communication can be conveyed. The problem at Level B (1120) is a semantic problem related to how accurately the conveyed symbols convey the desired meaning. The problem at Level C (1130) is an effectiveness problem related to how effectively the received meaning influences the operation in the desired manner.

[0164] Shannon's information theory focuses only on technical problems at Level A (1110). However, Weaver explains that Shannon's information theory is general enough to consider problems at Level B (1120) and Level C (1130) by adding semantic transmitters, semantic receivers, and semantic noise to Shannon's communication model.

[0165] Meanwhile, one of the various goals of 6G communication is to enable various new services that interconnect people and machines. Therefore, a semantic communication method that considers the semantic issues of Level B (1120) needs to be provided, going beyond the technical issues of Level A (1110) in communication systems. Semantic communication refers to the efficient transmission and reception of semantic information between a first device and a second device, respectively corresponding to the source and destination, using common background knowledge. Referring to the communication model of Figure 11, if the meaning of the intended message sent by the first device, which corresponds to the source, is accurately interpreted by the second device, which corresponds to the destination, it can be said that correct semantic communication has been performed.

[0166] For semantic communication, a source can generate a semantic representation based on given or collected raw data and transmit the generated semantic representation to a destination. The destination interprets and reasons the received semantic representation to align with the source's intent. Semantic communication requires an approach not from the perspective of reducing reconstruction errors that occur during the process of restoring the received semantic representation to the original raw data, but from the perspective of whether the downstream task performed by the destination can operate according to the source's intent using the received semantic representation. Thus, the semantic representation generated by the source and transmitted to the destination must be generated with consideration for the downstream task performed by the destination. That is, a task-oriented semantic communication system is needed that can generate semantic representations based on whether a task operation is performed well at the destination, and the task-oriented semantic communication system can preserve task-relevant information while introducing useful invariances to downstream tasks.

[0167] As described above, to perform semantic communication, a new layer for semantic communication, such as a semantic layer, can be defined that defines the overall operation of semantic expressions and messages. For example, a semantic layer can be located at each of the source and destination, reflecting a task-oriented semantic communication system. To perform communication between the semantic layers located at each of the source and destination, a protocol, which is a convention between the layers, and a series of operation processes need to be defined.

[0168] In semantic communication, messages received at the destination may contain errors due to channel noise between the source and destination. Errors can occur at either the technical or semantic level. The technical level seeks to preserve the syntactic integrity of the transmitted and received messages. Therefore, errors at the technical level can be defined as syntactic differences between the transmitted and received messages. In contrast, the semantic level does not seek to preserve the syntactic integrity of the transmitted and received messages, but rather seeks to infer similarities between the semantics of the transmitted and received messages. Therefore, errors at the semantic level can be defined as semantic differences between the transmitted and received messages.

[0169] An example of a semantic error is shown in Figure 12. Figure 12 illustrates an example of a semantic communication error. Figure 12 exemplifies a case where the source is a person and conveys a message through speech. Referring to Figure 12, the source attempts to transmit the message "copy machine" to the destination using sound. However, since the pronunciation of "p" is similar to the pronunciation of "ff," the destination may misinterpret the message as "coffee machine," contrary to the source's intention. This can be viewed as a semantic error that occurs when the semantic information transmitted from the source is misinterpreted at the destination. To prevent this type of semantic error, the source can transmit a different word, "Xerox," which has the same meaning, instead of the word "copy machine." In this case, the additionally received information and the previously received information are used together for semantic interpretation, so the semantic error can be prevented.

[0170] As mentioned above, semantic communication can result in semantic mismatches between the source and destination due to semantic errors. To ensure reliable semantic communication, the destination must detect semantic mismatches using semantic error checking, correct semantic errors, or request retransmission of semantic information from the source, thereby obtaining the exact semantic information intended by the source.

[0171] To detect semantic errors, attention techniques, first introduced in the field of neural machine translation (NMT), can be used. According to attention techniques, the decoder can refer back to the encoder's input sequence at each step of predicting the output word. In other words, the decoder does not consider all words in the input sequence with equal weight, but rather focuses on words highly related to the predicted word. This technique is called attention, and the information indicating the focused portion can be determined using an attention function.

[0172] Figure 13 illustrates an example of visualizing the association between words in a translation task using the attention technique. The attention technique can be applied to sequence modeling between input English words and output French words. Here, the sequence modeling between input English words and output French words can include weight values ​​expressed as a correlation matrix. In Figure 13, the closer the pixel color is to white, the higher the association between words. Referring to Figure 13, it can be confirmed that even if the order in which English and French sentences are written differs, the weight is higher for parts that are to be translated using words that have an association regardless of the position.

[0173] Transformers, designed to apply attention techniques to any network architecture and focus on all important parts of the input, regardless of input length, have been introduced. Using transformers, techniques capable of processing various data modalities, such as text, images, and graphs, are being studied. Attention techniques focus on parts of the input data that are highly relevant to the expected output. When applying attention techniques to semantic communication, the source can use attention techniques to generate and convey semantic representations to express its intended message. Based on the semantic representations conveyed using attention techniques, the destination can determine which parts of the original data to focus on to better interpret the meaning the source intended to convey. Therefore, utilizing attention techniques in semantic communication is desirable both for expressing the intended meaning and for interpreting the conveyed meaning.

[0174] Looking at the research conducted to date, it is common to configure systems using the transformer structure as is, perform training with a focus on reconstruction of input data, and use the trained model for inference. Furthermore, existing research does not focus on the attention map (e.g., attention matrix), which is the result of attention techniques. Therefore, treating the area indicated by the attention map as a semantic area corresponding to the meaning the source intends to convey, and utilizing the attention map for semantic communication expression and interpretation, is desirable from the perspective of setting and utilizing semantics.

[0175] Inductive bias refers to a set of assumptions a model makes to predict the output of an unknown input, i.e., additional assumptions made to account for contingencies to improve generalization performance. For example, CNNs are well-suited for image processing because they have inductive biases that assume dependencies (locality) between adjacent pixels and translation invariance, which recognize objects as the same even if their location changes. Another example is recurrent neural networks (RNNs), which use the concept of time. Similar to CNNs' inductive bias, RNNs have relational inductive biases that assume the characteristics of time series inputs and that the output is identical when inputs are received in the same order. In contrast, transformers perform self-attention operations that compute relationships between all elements of the input data, i.e., tokens, and can identify global regions regardless of the distance between data elements. When performing a self-attention operation, first, a matrix for attention weights can be determined as shown in [Mathematical Formula 1] below.

[0176]

[0177] In [Equation 1], is the attention weight matrix, Q is the query, K is the key, means the dimension of Q and K. Here, Q and K are matrices obtained through embedding so that the computer can recognize the input data in units of set tokens (e.g., dimension is d). embed ) and the learned weight matrix W Q and W KThese are matrices obtained by multiplying. For example, if the input data is a sentence, the tokens can be set in units of words. The query refers to an information vector that attempts to find a related part in the input sequence, and the key refers to a vector that is compared with the query to determine the degree of relationship relevance.

[0178] In [Equation 1], is used as a scaling factor to perform scaling so that the size of the value after matrix multiplication of Q and K does not become too large. In addition, QK T The size of is n×n if the number of input tokens that the model with the attention technique can process at one time is n. Therefore, since it can be seen that the global region is identified by performing an operation to check all correlations for the same input data, the inductive bias of the transformer is relatively lacking compared to CNN and RNN due to this operation. In order to overcome this lack of inductive bias, a method is mainly adopted to improve the performance of the attention technique by using a large dataset for learning for transformer-based models.

[0179] Therefore, when applying a transformer based on the attention technique to semantic communication, the size of the dataset will increase to improve the performance of the model due to the lack of inductive bias, which may incur overhead from a communication perspective as frequent communication between the source and the destination, i.e., forward propagation operation for learning and back-propagation operation for gradient transfer, is frequently required.

[0180] Additionally, the transformer model can have a multi-head attention structure that performs the aforementioned self-attention operation in parallel. Here, the parallel operation can be performed h times depending on the number of heads set, h. Here, since each head attends to a different part of the input data, the model can handle more complex relationships between input tokens, which can have the advantage of representing the input data with more information. However, each head focuses on a different area of ​​the data, which can be interpreted as each head paying attention to a different area of ​​the data. Multi-head self-attention can be interpreted as paying attention to most of the entire area of ​​the input data. Figure 14 illustrates an example of the results of visualizing the attention map for multi-heads for image-type input data. Referring to Figure 14, it can be confirmed that as the layer expressed as a block becomes deeper, most of the global area is captured.

[0181] When introducing the transformer technique to semantic communication, if the attention map obtained through the attention technique in the final layer of the transformer model is used, it becomes impossible to know that the source expected from the attention technique only attends to the intended desired region, and thus it may become difficult to perform semantic error detection and correction for the semantic region using the attention map.

[0182] Therefore, to address the aforementioned issues, the present disclosure proposes a technique for additionally generating priors, which correspond to structural information about input data from the source—information not addressed by existing techniques—and applying an attention technique that utilizes the priors together. Furthermore, the present disclosure proposes specific techniques for performing the learning and inference operations necessary for semantic error detection and correction based on the attention map obtained by utilizing the priors.

[0183]

[0184] The present disclosure proposes a technique for performing semantic error detection and correction using an attention map generated through an attention technique that considers input data and prior information related to the input data together to express the meaning intended to be conveyed by a source in a system supporting semantic communication. The proposed technique addresses a semantic level corresponding to Level B in FIG. 11 and can be used for two-way communication. Prior information related to input data from the source has a similar role to "xerox" illustrated in FIG. 12 and corresponds to additional knowledge for expressing the meaning intended to be conveyed by the source. Accordingly, the destination can interpret the source's intent using the message and prior information corresponding to the received information, as illustrated in FIG. 12.

[0185] In a system according to embodiments of the present disclosure, learning for semantic communication is performed. According to one embodiment, the method of performing learning can be broadly divided into two types. For example, the source and destination can perform learning for extracting attention maps and performing semantic error detection and correction based on the received and generated attention maps, targeting input data and a model for an encoder that extracts prior information about the input data and applies an attention technique, a model for a decoder that extracts reconstructed data at a semantic level, and a model for a downstream task located at the destination. As another example, if an existing transformer-based encoder model exists, learning can be performed by performing initialization using weights for the model, extracting an attention map using prior information in the same manner as in the above-described example, and performing fine-tuning on the encoder that performs learning related to semantic error detection and correction. Through fine-tuning, the weights of not only the model for the encoder, but also the models for the decoder and downstream tasks can be initialized using the weights of previously used models.

[0186] Once the learning operation is complete, encoders located at the source and destination can generate prior information based on the input data, and an attention map can be created based on this prior information. Accordingly, semantic error detection and correction operations can be performed based on the attention map. Furthermore, with regard to the meaning intended to be conveyed across the entire data, the present disclosure proposes a system that sets the data region to be focused on as a semantic element and uses the semantic information to enable downstream tasks located at the destination to operate according to the desired intent.

[0187] The structure of the source related to the semantic representation, attention map, and prior information for extracting semantic regions by utilizing prior information together with the service operation of each task at the source and destination is as shown in Fig. 15 below, and the structure of the destination is as shown in Figs. 16a and 16b below.

[0188] FIG. 15 illustrates the functional structure of a source for performing attention map-based semantic error detection and correction considering priors according to one embodiment of the present disclosure. Referring to FIG. 15, the source includes a semantic data retransmission count checker (1502) that checks a retransmission count value, an attention map receipt checker (1504) that checks the reception of an attention map, an encoder (1506) that generates a semantic expression and prior information, an attention map generator (1508) that generates an attention map for a semantic expression, a semantic packet generator (1510) that generates at least one packet for transmitting the semantic expression, prior information, and attention map, a data saving block (1512) that stores information required for semantic error correction, a residual area data extractor (1514) that extracts residual area data from original data, and a system failure operator (system failure operator) that determines the failure of semantic error correction. The encoder (1506) may be composed of multiple layers to which an attention technique is applied, and may generate prior information based on input data during an encoding operation. The attention matrix that may be obtained during the encoding operation may include a matrix for attention scores before applying the softmax operation in [Mathematical Formula 1].Furthermore, the attention value matrix may be included, which is ultimately obtained by multiplying V, which is a value representing a vector used to determine weights as information on an input sequence corresponding to a specific key, after calculating the AttentionWeight of [Mathematical Formula 1], in addition to the matrix for the attention score before applying the softmax operation. At this time, the value is also a weight matrix learned in a matrix in which embedding is performed on the input data for each set token unit. can be obtained by multiplying, and the dimension of the matrix is am.

[0189] FIGS. 16A and 16B illustrate the functional structure of a destination for performing attention map-based semantic error detection and correction considering priors according to one embodiment of the present disclosure. Referring to FIGS. 16A and 16B, the destination comprises a semantic packet decomposer (1602) that obtains data (e.g., semantic expression, prior information) from at least one semantic packet received from a source, a decoder (1604) that reconstructs the semantic expression obtained from at least one semantic packet at a semantic level, a reconstructed data type checker (1606) that verifies the type of data reconstructed at a semantic level (e.g., original data, residual region data), a checker for performing semantic error checking (1608) that determines whether to perform semantic error checking based on a setting, a data type configurator (1610) that sets the type of data and whether to correct errors, an encoder (1612) that generates a semantic expression and an attention matrix for error checking and / or error correction, An attention map generator (1614) for generating an attention map based on an attention matrix, a semantic error checker (1616) for determining a semantic error based on information related to an attention map provided from a data storage block (1620) for storing information related to an attention map received from a source or data for checking and correcting semantic errors and attention map information generated by the attention map generator (1614), an early stopping checker (1618) for determining whether to stop early in response to the occurrence of a semantic error, a data storage block (1620) for storing data for checking and correcting errors (e.g., received prior information, received attention map, generated attention map, etc.),An attention map based packet generator (1622) that generates at least one packet including information related to an attention map fed back to a source, a system failure operator (1624) that performs an action according to a failure of semantic error correction, a reconstructed data type checker (1626) that checks the type of data reconstructed at a semantic level (e.g., whether it is residual region data), a data compositor (1628) that synthesizes residual region data reconstructed at a semantic level into previously stored data, and a plurality of downstream task blocks (163-1 to 1630-N) that perform tasks according to settings. Based on a structure such as FIG. 16, the prior information received from the source can be used for the operation of the encoder (1612) of the destination and for calculating an attention matrix for input data. Additionally, downstream tasks can perform actions according to the purpose of the task by using the semantic representation obtained as a result of the execution operation of the encoder located at the destination as input.

[0190]

[0191] Figure 17 illustrates an example of a procedure for performing semantic communication by a first device according to one embodiment of the present disclosure. Figure 17 illustrates a method performed by the first device. The first device acts as a source or destination for performing semantic communication, and may be a terminal, a base station, or any entity defined in a wireless communication system.

[0192] Referring to FIG. 17, at step S1701, the first device establishes a connection with the second device. To this end, the first device may receive a synchronization signal from the second device, acquire synchronization with the second device based on the synchronization signal, and perform signaling to establish the connection. Alternatively, the first device may transmit a synchronization signal and perform signaling to establish the connection.

[0193] In step S1703, the first device transmits capability information to the second device. In other words, the first device may transmit a message including capability information to the second device. Prior to this, although not illustrated in FIG. 17, the first device may receive a request for capability information from the second device. Here, the capability information may include information about functional and hardware capabilities related to communication of the first device. According to one embodiment, the capability information may include capability information related to semantic communication, particularly, semantic error detection and correction. For example, the capability information related to semantic error detection and correction may include at least one of information indicating that operations related to an attention map using a priori are possible and information indicating that detection and correction of semantic errors are possible.

[0194] In step S1705, the first device receives configuration information related to communication from the second device. The configuration information may be received via at least one message. The configuration information may indicate resources for communication, data processing related to communication, values ​​of variables related to communication, etc. Accordingly, the first device may store the configuration information and set variables, parameters, etc. related to a receiver or transmitter for communication. According to various embodiments, the configuration information may include various information, variables, parameters, etc. for performing semantic error detection and correction based on an attention map using a prior.

[0195] In step S1707, the first device performs semantic communication based on the configuration information. In addition, the first device may support or perform semantic error detection and correction based on the configuration information. If the first device is a source, the first device may transmit at least one semantic expression for multi-tasks, and further, according to one embodiment, may transmit attention map information corresponding to the semantic expression using prior information, and, if necessary, may perform retransmission based on a residual region used for semantic error correction. If the first device is a destination, the first device may perform data restoration at a semantic level from the received at least one semantic expression for multi-tasks, detect semantic errors in the data restored at the semantic level, and, depending on whether or not there is a semantic error, may request retransmission based on a residual region used for semantic error correction so as to perform a semantic error correction operation.

[0196]

[0197] Figure 18 illustrates an example of a procedure for performing semantic communication by a second device according to one embodiment of the present disclosure. Figure 18 illustrates a method performed by the second device. The second device acts as a destination or source for performing semantic communication, and may be a base station, a terminal, or any entity defined in a wireless communication system.

[0198] Referring to FIG. 18, at step S1801, the second device establishes a connection with the first device. To this end, the second device may transmit a synchronization signal and perform signaling to establish the connection. Alternatively, the second device may receive a synchronization signal from the first device, acquire synchronization with the first device based on the synchronization signal, and perform signaling to establish the connection.

[0199] In step S1803, the second device receives capability information from the first device. In other words, the second device may receive a message including capability information from the first device. Prior to this, although not illustrated in FIG. 18, the second device may transmit a request for capability information to the first device. Here, the capability information may include information about functional and hardware capabilities related to communication of the second device. According to one embodiment, the capability information may include capability information related to semantic communication, particularly, semantic error detection and correction. For example, the capability information related to semantic error detection and correction may include at least one of information indicating that operations related to an attention map using a priori are possible and information indicating that detection and correction of semantic errors are possible.

[0200] In step S1805, the second device transmits configuration information related to communication to the first device. The configuration information may be transmitted via at least one message. The configuration information may indicate resources for communication, data processing related to communication, values ​​of variables related to communication, etc. Accordingly, the second device may store the configuration information and set variables, parameters, etc. related to a receiver or transmitter for communication. According to various embodiments, the configuration information may include various information, variables, parameters, etc. for performing semantic error detection and correction based on an attention map using a prior.

[0201] In step S1807, the second device performs semantic communication based on the configuration information. In addition, the second device may perform or support semantic error detection and correction based on the configuration information. If the second device is a destination, the second device may perform data restoration at the semantic level from at least one semantic expression for the received multi-task, detect semantic errors in the data restored at the semantic level, and, depending on whether an error exists, request retransmission based on the residual region used for semantic error correction so as to perform a semantic error correction operation. If the second device is a source, the second device may transmit at least one semantic expression for the multi-task, and further, according to one embodiment, transmit attention map information corresponding to the semantic expression using prior information, and, if necessary, perform retransmission based on the residual region used for semantic error correction.

[0202]

[0203] Figure 19 illustrates an example of a procedure for performing semantic communication using prior information of a first device according to one embodiment of the present disclosure. Figure 19 illustrates a method performed by the first device. The first device acts as a source or destination for performing semantic communication, and may be a terminal, a base station, or any entity defined in a wireless communication system.

[0204] Referring to FIG. 19, in step S1901, the first device determines a prior for the input data. For example, the prior may include information related to the distribution of the input data. As another example, the prior may include an intermediate result generated during the computational process for generating an attention map.

[0205] In step S1903, the first device generates an attention map using a prior. The attention map is information indicating a focused area in the input data in the process of generating a semantic representation to indicate the intention that the first device wants to convey. According to various embodiments, the first device may generate a prior-based attention map by fusing an attention map based on self-attention and a prior. For example, the first device may generate a prior-based attention map by performing a weighted sum operation of an attention matrix generated by a first layer and an attention logit matrix generated by a second layer, which is a layer preceding the first layer, among layers that perform operations related to an attention map included in a network for generating an attention map. Here, the first layer may have a greater depth than the second layer. As another example, the first device can generate a prior-based attention map by performing a softmax operation after adding information related to the distribution of input data and information related to the relative position of each token to the target of a softmax operation for self-attention.

[0206] In step S1905, the first device transmits a semantic packet including a prior-based attention map. In addition to the prior-based attention map, the semantic packet may further include one of a semantic representation or prior information. Then, based on the feedback information, the second device may further transmit a semantic packet including at least one of a semantic representation for the remaining region, a prior-based attention map, or prior information.

[0207]

[0208] Figure 20 illustrates an example of a procedure for performing semantic communication using prior information of a second device according to one embodiment of the present disclosure. Figure 20 illustrates a method performed by the second device. The second device acts as a destination or source for performing semantic communication, and may be a base station, a terminal, or any entity defined in a wireless communication system.

[0209] Referring to FIG. 20, in step S2001, the second device determines a prior for the restored data. Here, the restored data includes data restored at a semantic level. Although not illustrated in FIG. 20, the second device may receive a semantic packet including a semantic expression from the first device, and decode the semantic expression obtained from the semantic packet to obtain data restored at a semantic level. In addition, the second device may generate a semantic expression by encoding the data restored at a semantic level, and may obtain a prior in the process. For example, the prior may include information related to the distribution of input data. As another example, the prior may include an intermediate result generated during a computational process for generating an attention map.

[0210] In step S2003, the second device generates an attention map using a priori. The attention map is information indicating a focused area from the input data of the second device to check whether it matches the focused area that matches the intention that the first device wants to convey. Here, the input data of the second device may include data restored to a semantic level. It is information indicating a focused area in the process of generating a semantic representation. According to various embodiments, the second device may generate a priori-based attention map by fusing an attention map based on self-attention and a priori. For example, the second device may generate a priori-based attention map by performing a weighted sum operation of an attention matrix generated by a first layer and an attention logit matrix generated by a second layer, which is a layer preceding the first layer, among layers that perform operations related to an attention map included in a network for generating an attention map. Here, the first layer may have a greater depth than the second layer. As another example, the second device can generate a prior-based attention map by performing a softmax operation after adding information related to the distribution of input data and information related to the relative position of each token to the target of the softmax operation for self-attention.

[0211] In step S2005, the second device detects semantic errors based on the prior-based attention map. Although not illustrated in FIG. 20, the second device may receive a semantic packet including a prior-based attention map from the first device, and detect semantic errors by comparing the prior-based attention map obtained from the semantic packet with the prior-based attention map generated in step S2003.

[0212] The embodiment described with reference to FIG. 20 can be understood as a procedure in which the second device performs inference. When training is performed, in addition to the procedure illustrated in FIG. 20, the following operations may be further performed. According to one embodiment, the second device may generate a priori based on input data, as in step S2003, and may also obtain a priori generated from the input data of the first device and then transmitted, and perform learning so as to reduce the difference between the two priori, in other words, so as to generate a priori similar to the priori received from the first device.

[0213]

[0214] Figure 21 illustrates an example of a procedure supporting attention map-based semantic error detection and correction considering priors according to one embodiment of the present disclosure. Figure 21 illustrates a method performed by a device operating as a source. At least one of the operations illustrated in Figure 21 may be omitted depending on circumstances and / or settings.

[0215] Referring to Figure 21, in step S2101, the source checks whether there is data to be transmitted. If there is data to be transmitted, in step S2103, the source sets a variable indicating the number of times the semantic data is retransmitted. Check whether is less than the threshold. Here, the threshold is It could be. If the threshold is greater than the threshold, in step S2105, the source performs a system failure action. For example, the source may perform a reconfiguration for semantic error correction in response to the system failure or transmit information related to new data. As another example, the source may reset the semantic data retransmission count to 0 and flush the stored attention map information.

[0216] If the value is below the threshold, in step S2105, the source checks whether attention map information has been received. Here, attention map information is transmitted by the destination to request retransmission of data for correction of semantic errors. If attention map information has not been received, in step S2107, the source determines whether the variable is set as new transmission data. If attention map information is received, in step S2109, the source is a variable Increases by 1. Then, in step S2111, the source extracts residual region data based on attention map information. Here, the attention map information may include at least one of attention map information received from the destination or attention map information extracted based on original data. In step S2113, the source extracts a variable Set as residual area data.

[0217] Afterwards, in step S2115, the source performs encoding using prior information. At this time, the source can use an attention mechanism, and the prior information can be generated from data transmitted to the encoder. In step S2117, the source generates an attention map reflecting the prior information. Then, in step S2119, the source determines whether bitwise operation is performed. If bitwise operation is performed, in step S2121, the source converts each attention map into a bitwise attention map. For this purpose, a threshold This can be utilized. Then, in step S2123, the source transmits at least one semantic packet. The semantic packet may include at least one of a semantic expression, an attention map, and prior information. Then, in step S2125, the source stores the attention map transmitted to the destination.

[0218]

[0219] Figures 22a and 22b illustrate examples of a procedure for performing attention map-based semantic error detection and correction considering priors according to one embodiment of the present disclosure. Figures 22a and 22b illustrate a method performed by a device operating as a destination. At least one of the operations illustrated in Figures 22a and 22b may be omitted depending on circumstances and / or settings.

[0220] Referring to FIGS. 22A and 22B , in step S2201, the destination may receive at least one semantic packet from the source. Here, the semantic packet may include at least one semantic expression, an attention map, or prior information. In step S2203, the destination decomposes the at least one semantic packet. Through this, the destination may obtain at least one semantic expression, an attention map, or prior information. In step S2205, the destination performs decoding on the at least one semantic expression. Through this, the destination may obtain data restored to a semantic level.

[0221] Afterwards, in step S2207, the destination verifies the type of data restored to the semantic level. If the data restored to the semantic level is data restored to the semantic level based on the original data, in step S2209, the destination checks the type of data restored to the semantic level. is set as reconstructed data at the semantic level based on the original data. If the reconstructed data at the semantic level is data reconstructed at the semantic level based on the residual region for semantic error correction, in step S2211, the destination determines whether to perform semantic error checking. If semantic error checking is performed, in step S2213, the destination determines whether to perform a variable The destination sets the data restored to the semantic level based on the residual region. If semantic error checking is not performed, in step S2215, the destination composites the data. That is, the destination performs data synthesis as a process of semantic error correction. For example, the destination can synthesize data restored to the semantic level based on the residual region and data restored to the stored semantic level (e.g., data that is the target of semantic error correction). Then, in step S2217, the destination sets the variable Set the data restored to the semantic level based on the synthesized data.

[0222] Subsequently, in step S2219, the destination performs encoding using the prior information. At this time, the destination may utilize prior information received from the source, prior information obtained from data restored at the semantic level by the destination, and an attention mechanism. Additionally, if training is required, the destination may perform training to reduce the difference between the obtained prior information and the received prior information, i.e., to generate priors similar to the priors received from the source.

[0223] Next, in step S2221, the destination generates an attention map reflecting the prior information. In step S2223, the destination determines whether bitwise operations are performed on the attention map. If bitwise operations are performed, in step S2225, the destination generates a bitwise attention map. For this purpose, a threshold value is used. This can be used. And, in step S2227, the destination attempts to detect semantic errors. Semantic errors are detected when the threshold can be detected using . In step S2229, the destination checks whether a semantic error has occurred.

[0224] If no semantic error is detected, in step S2231, the destination verifies the input data type of the encoder. If the data type is data restored to a semantic level based on the residual region, the destination proceeds to step S2215 and performs data synthesis, which is the semantic error correction process described above. On the other hand, if the data type is not data restored to a semantic level based on the residual region, that is, if the data type is data restored to a semantic level based on the original data or synthesized data of a semantic level that has undergone a semantic error correction process, in step S2233, the destination performs a downstream task.

[0225] On the other hand, if a semantic error is detected, in step S2235, the destination determines whether to perform an early termination. If it is determined that an early termination will be performed, in step S2237, the destination performs a system failure operation. If it is determined that an early termination will not be performed, in step S2239, the destination stores data (e.g., data restored to the semantic level, attention map information based on the data restored to the semantic level, attention map information based on the original data, etc.). Then, in step S2241, the destination transmits information related to the attention map. That is, by transmitting information related to the attention map, the destination requests transmission of a semantic packet including additional information for correcting a semantic error.

[0226]

[0227] FIG. 23 illustrates an example of an initial setup procedure for semantic communication according to one embodiment of the present disclosure. FIG. 23 illustrates signaling for setup related to semantic communication between a first device (2310) and a second device (2320). Here, the first device (2310) may be one of a source and a destination, and the second device (2320) may be the other of the source and the destination. Furthermore, the first device (2310) may be a terminal and the second device (2320) may be a base station, or the first device (2310) may be a base station and the second device (2320) may be a terminal, or both the first device (2310) and the second device (2320) may be terminals.

[0228] Referring to FIG. 23, in step S2301, the first device (2310) and the second device (2320) perform synchronization and connection establishment procedures. To this end, the first device (2310) and the second device (2320) may perform synchronization using at least one synchronization signal and perform signaling for connection establishment. If necessary, the first device (2310) and the second device (2320) may perform a random access procedure for initial access. Specifically, the first device (2310) may transmit a random access preamble, receive a random access response message from the second device (2320), and transmit an uplink message through resources allocated by the random access response message.

[0229] In step S2303, the second device (2320) transmits a capability information request message to the first device (2310). For example, the second device (2320) transmits a downlink-dedicated control channel (DL-DCCH) message including a UE capability inquiry indicator to the first device (2310). In particular, according to one embodiment, the capability information request message includes information inquiring whether semantic communication can be performed.

[0230] In step S2305, the first device (2310) transmits a capability information message to the second device (2320). For example, the first device (2310) transmits an uplink-dedicated control channel (UL-DCCH) message including UE capability information. In particular, according to one embodiment, the capability information message includes capability information related to semantic communication. According to various embodiments, the capability information may include information on whether the first device (2310) has a semantic communication performance capability. In addition, the capability information may further include information on the type of raw data that can be generated, collected, or processed by the first device (2310), the computational capability of the first device (2310), and the like. According to various embodiments of the present disclosure, the capability information message may include information indicating that an operation related to an attention map using a priori is possible, and information indicating that detection and correction of semantic errors are possible.

[0231] In step S2307, the second device (2320) may determine whether to perform semantic communication. For example, the second device (2320) may recognize that the first device (2310) has the capability to perform semantic communication based on capability information received from the first device (2310). In addition, the second device (2320) may further consider other acquired information to determine whether to perform semantic communication. Specifically, the second device (2320) may determine whether the semantic communication-related capability of the first device (2310) satisfies the requirements for semantic communication. In the example of FIG. 23, it is determined whether to perform semantic communication.

[0232] In step S2309, the second device (2320) transmits information related to semantic communication to the first device (2310). Here, the configuration information may include information instructing the performance of semantic communication and configuration information for performing semantic communication. The information related to semantic communication may be transmitted via at least one of a downlink control information (DCI), a medium access control (MAC) control element (CE), or a radio resource control (RRC) message. Specifically, the information instructing the performance of semantic communication may be transmitted via upper layer signaling, and at least a portion of the configuration information may be transmitted via DCI or MAC CE.

[0233] In step S2311, the first device (2310) stores information related to semantic communication. According to one embodiment, the configuration information may indicate parameters, variables, etc. related to error detection and correction of semantic communication. Specifically, the configuration information may include at least one of information related to whether fine-tuning is performed, information related to an encoder / decoder model, information related to a technique for generating a priori, information related to a technique for transmitting a priori, information related to whether a bit-wise attention map is set, information related to a threshold related to a bit-wise attention map, information related to a threshold related to error detection, information related to a maximum semantic data retransmission count, information related to an early termination threshold, and information related to whether semantic error detection is performed on remaining region-based data.

[0234] Thereafter, although not illustrated in FIG. 23, the first device (2310) and the second device (2320) may perform semantic communication according to various embodiments of the present disclosure. At this time, the first device (2310) and the second device (2320) may perform semantic error detection and semantic error correction based on the information transmitted in step S2309 and stored in step S2311.

[0235] As shown in Fig. 23, an initial setup procedure for semantic communication using prior information can be performed. In this case, when learning through fine-tuning is performed, initialization can be performed using the weights of the pre-learned encoder and decoder models. Otherwise, the encoder and decoder located at the source and destination can randomly initialize the weights of the model and perform learning. The technique for generating and utilizing priors and how to transfer the generated prior information from the source to the destination can be set through the initialization operation and applied to the encoding operation located at the source and destination. Depending on whether a bit-wise attention map is generated, the attention matrix obtained during the encoding operation can be converted into a bit-wise form or used as is. In addition, the threshold related to the bit-wise attention map is as described above. It can be. In addition, the threshold related to the semantic error rate, which is the criterion for detecting semantic errors, is It can be, and the maximum semantic data retransmission counter for mantic error correction operations is The early termination threshold, which is the error rate associated with the criteria for performing semantic error correction operations at the destination, may be Finally, depending on whether semantic error detection is performed on the remaining region-based data extracted during the semantic error correction operation, it may be determined whether semantic error detection and correction operations are performed on the corresponding data at the destination. According to various embodiments, some of the steps described with reference to FIG. 23 may be omitted depending on the circumstances and / or settings.

[0236]

[0237] Hereinafter, the present disclosure describes a specific example of the above-described semantic error detection and correction. As an example of a procedure for performing semantic error detection and correction according to one embodiment of the present disclosure, an example of an operation of generating an attention map for measuring semantic errors together with prior information and transmitting a semantic expression and prior information including the content intended by the source to the destination is as follows. In the following description, it is assumed that the encoders included in the source and destination use the weights of a previously learned transformer-based encoder model for initialization and then perform fine-tuning based on the weights.

[0238] 1) At the beginning of semantic communication, there is no attention map received from the source to the destination. In addition, the value of the counter for semantic error correction is a set number of times to prevent indiscriminate retransmission for semantic error correction. is smaller than . That is, for the initial transmission, the counter for semantic error correction is 0. Therefore, if the input data (original data) exists, the destination checks whether the received attention map exists. Similarly, at the beginning of semantic communication, since the received attention map does not exist, the destination performs an encoding operation.

[0239] 2) The source performs an encoding operation on the input data. For example, when calculating the attention map as in [Mathematical Formula 1], the source obtains a new, final attention map by fusing the prior information on the input data. At this time, the method of extracting the prior information and fusing it with the existing attention map information can be defined in various ways. The fusion operation can be conceptually represented as in Fig. 24. Fig. 24 illustrates an example of a final attention map generated by utilizing the prior information. Referring to Fig. 24, it may be considered to use some of the information related to the operation for the attention map as prior information. More specifically, when the layers performing the operation related to attention are connected through residual connections, and the current layer performs the operation related to the attention map, the current layer can receive the intermediate result obtained from the operation for the attention map in the previous layer as a prior through the residual connections and perform an operation that uses the received prior information together, i.e., the fusion operation. At this time, the prior information can be configured in a form that adds an inductive bias to the CNN. The aforementioned operation is expressed mathematically as [Mathematical Equation 2] below.

[0240]

[0241] In [Equation 2], is the input data, silver Attention logit matrix calculated in the layer before the th layer, is the current before the softmax operation The logit matrix operated on the input data in the th layer, is a 2D-convolution layer with ReLU activation applied, and refers to weight factors having values ​​in the range of 0 to 1. Here, if it is the first layer that performs the operation, does not exist. is identical to the matrix for the attention score, which is the value before performing the softmax operation in [Mathematical Formula 1]. For the operation, it can be implemented by adding an additional network to the 2D-convolutional layer.

[0242] The layer structure for performing [Mathematical Formula 2] is as shown in FIG. 25. FIG. 25 illustrates an example of a layer structure for generating a prior-based attention map using residual connections and convolution operations according to one embodiment of the present disclosure. In FIG. 25, blocks (2502-1 and 2502-2) are layers in an encoder to which the attention technique is applied. Referring to FIG. 25, the results calculated through the convolution operation and the attention technique in the previous layer are transmitted as prior information through residual connections between layers.

[0243] In the above example, the prior information may include information related to the attention map computed in the previous layer of the layer to which the current attention technique is applied. Information related to the attention map computed in the previous layer may exist in each layer except the first layer, and may be present in the current layer. Based on the th layer, [Mathematical Formula 2] can be expressed as . Similarly, the following The fryer information to be transmitted to the layer is [Mathematical Formula 2] can be expressed as . Therefore, the number of priorities that can be passed from the source to the destination can be {the number of layers that constitute the encoder - 1}.

[0244] Another example of using a fryer is fryer information. If the attention weight operation is generated based on the distribution of input data and the fusion operation is defined as addition, the attention weight operation modified from [Mathematical Formula 1] can be defined as in [Mathematical Formula 3] below.

[0245]

[0246] In [Equation 3], the prior information It is constructed under the assumption that follows a Gaussian distribution, which means that when calculating the attention technique, the influence of tokens close to the token to be focused on may be greater than that of tokens that are relatively far away. In [Mathematical Formula 3], It can be determined based on one of the embedding techniques related to the relative position indicating the relative position of each token.

[0247] In [Equation 3] Constituting which represents the central position and tokens If defined as the tightness between can be determined as shown in [Mathematical Formula 4] below.

[0248]

[0249] In [Equation 4] is from the entire dataset This refers to the window size for the second query. The window size is a factor that determines how far to the left and right from the center position the influence of tokens will be reflected.

[0250] Prior applied to self-attention operation of [Equation 3] for and can be determined as in [Mathematical Formula 5].

[0251]

[0252] In [Equation 5] is the length of the input data, i.e. the number of input tokens, is the central location of the distribution related to the prior information generated based on the distribution of the input data, refers to the window size, which is a factor that determines how far away from the center position the influence of tokens will be reflected. Is and The value of 0 to It is used to adjust the value within the range.

[0253] Here, and can be determined through learning. can be set by learning as in [Mathematical Formula 6] below, and has a scalar value.

[0254]

[0255] In [Mathematical Formula 6], the center of the sequence provided as input to the model is Assuming that, the query The second row is can be expressed as, is a weight matrix corresponding to the learned parameters related to the central position, can be composed of a learnable feed-forward network that performs linear projection. can be set by learning as in [Mathematical Formula 7] below, and has a scalar value.

[0256]

[0257] In [Equation 7], can be defined with the same parameters as [Mathematical Formula 6], can be composed of a learnable feed-forward network that performs linear projection. Sharing can be understood as the central location and window size interdependently specifying the location for the prior information.

[0258] To summarize, by [Equation 4] and [Equation 7] Prioritize based on the center position and window size set for the second query. Element of can be calculated, and each of the remaining elements can be determined to a different value depending on the method of fusion. In the example, since fusion is expressed as addition as in [Mathematical Formula 3], it cannot be calculated by the center position and window size. Elements of can be set to 0.

[0259] [Mathematical expression 5] can be modified as shown in [Mathematical expression 8] below to apply it to a multi-head.

[0260]

[0261] In [Equation 8] and is the total number of heads When it's a dog The center position and window size of the th attention head can be learned as separate parameters for each head, so that It can be understood as allocating .

[0262] Furthermore, since [Equation 3] is used in each of the multiple layers that constitute the encoder, it can be determined whether or not to use the prior information in all layers. Previous research has shown that some lower layers (e.g., layers close to the input layer) obtain relatively more information contained in the prior compared to upper layers, and it has been revealed that applying the prior information to all layers actually results in a decrease in performance. Based on these results, a model can be designed so that only some lower layers use the prior information to compute the attention map. The number of lower layers that use the prior information may vary depending on the number of applied layers and the performance of the downstream task. For example, in a vanilla transformer, when an attention map operation based on the prior information is performed, only the lower three layers of the six layers that constitute the encoder can perform the attention map-related operation using the prior information.

[0263] In the above example, the extracted fryer information is Passing matrix information about the center position of each head and window size You can pass two values ​​as a set.

[0264] If a separate attention operation is performed, Instead of performing a series of rounds and concatenating the results, we replace it with a single self-attention calculation. , , The size of at If it is generated and operated as , for each layer There may be sets of the center positions and window sizes of the dog, and these sets can be passed to the destination. In this case, silver , and the dimension of the embedding matrix of token units set as input data in [Mathematical Formula 1], i.e., the input shape of the encoder. The same value can be used.

[0265] In the example of the process of extracting prior information described above, two types of prior extraction techniques were described. However, other prior extraction techniques may be applied in addition to the aforementioned prior information extraction techniques.

[0266] 3) The source encoder extracts the attention matrix, which is a matrix corresponding to the attention score, in the final layer and processes the attention matrix depending on whether bitwise operation is performed. At this time, the attention matrix may include the value before performing the softmax operation in [Mathematical Formula 1] or [Mathematical Formula 3]. In case of performing bitwise operation, the source encoder performs a normalization operation and the threshold value By performing bitwise operations using , a bitwise attention map can be obtained. Here, whether to create a bitwise attention map and can be set during the initial setup process. If no bitwise operations are performed, the source encoder can use the attention matrix as an attention map.

[0267] 4) The source generates semantic packet(s) based on the semantic expression(s), prior information, and attention map information obtained through steps 1) to 3), and transmits the semantic packet(s) to the destination. Then, the source stores the attention map information.

[0268] 5) The destination receives semantic packet(s) and decomposes the packets to obtain an attention map, prior information, and semantic representation(s). Then, the destination uses the obtained semantic representation(s) as input to the decoder to perform restoration at the semantic level. Here, restoration at the semantic level means the operation of restoring data to a level where the destination can obtain an output corresponding to the intended operation of the source by performing the operation of the downstream task normally or in accordance with the intention.

[0269] 6) The destination performs an encoding operation using the data restored to the semantic level and the received prior information. The encoding operation of the destination can be performed similarly to the operation described in 2). At this time, the encoder can be trained so that the prior information extracted during the encoding operation is similar to the prior information transmitted from the source. According to one embodiment, when the prior utilization technique using residual connections is used in 2), the encoder of the destination performs an encoding operation. Information and sources are communicated In order for the information to be similar to each other, [Equation 2] A 2D-convolutional layer for the operation can be trained. In another embodiment, when transmitting information related to the distribution of the input data in 2), the destination prior The center location and window size of the distribution related to the prior information, which are the results of the operations related to [Mathematical Equation 4] and [Mathematical Equation 7] for the multi-head required to obtain and The value for is passed from the source So that the values ​​in the set are extracted similarly to those in the destination. and The networks used to extract can be trained. By the operation of the destination encoder, semantic representation(s) can be obtained, and the obtained information can be used using semantic error checking and can be used as input for downstream tasks.

[0270] 7) Similar to 3), the destination encoder can extract an attention matrix corresponding to the attention score in the final layer and generate a bit-wise attention map depending on whether bit-wise operations are performed. Alternatively, the destination encoder can use the attention matrix as an attention map.

[0271] 8) The destination measures the semantic error rate by comparing the attention map received from the source with the attention map generated at the destination. Here, when measuring the semantic level error, the destination may not consider the semantic level error of the area other than the area on which it wants to focus. Therefore, the destination can identify the area on which it wants to focus using the attention map that considers the prior information, and measure the semantic error rate using the attention maps obtained from the source and the destination as in [Mathematical Formula 9] below.

[0272]

[0273] In [Equation 9], is the semantic error rate, It means the ratio of the area of ​​the attention map using a prior based on data restored to the semantic level among the areas of the attention map using a prior based on the data to be transmitted.

[0274] An example of [Mathematical Formula 9] is shown in FIG. 26 below. FIG. 26 illustrates examples of attention maps compared to determine a semantic error rate according to an embodiment of the present disclosure. FIG. 26 illustrates examples of attention maps that have performed bit-by-bit operations. Referring to FIG. 26, an attention map (2614) generated from original data (2612) and an attention map (2624) generated from data (2622) restored to a semantic level can be compared to determine a semantic error rate. That is, through a data region obtained according to an attention operation using prior information, attention maps are obtained that specify regions that allow the destination to obtain an output corresponding to the operation intended by the source by performing the operation of the downstream task normally or in accordance with the intention, and the regions are compared with each other. Here, the regions can be treated as semantic regions.

[0275] In Fig. 26, the attention map (2624) extracted from the destination is generated using priors extracted using a network trained to extract prior information similar to the prior information transmitted from the source, as described in 6). In operations related to the attention technique, the attention map (2624) can help generate an attention map similar to the attention map transmitted from the source using the prior information extracted from the destination. Accordingly, the destination can extract an attention map similar to the attention map region generated and transmitted from the source.

[0276] 9) The semantic error rate measured in 8) above is the threshold for determining whether there is a semantic error. If the semantic representation, which is the final output of the encoder located at the destination, is used as input for the downstream task, and an operation appropriate to the task's purpose is performed. Then, the loss for the task execution can be calculated. If the measured semantic error rate is below the threshold, If it is higher, a procedure for correcting semantic errors is performed.

[0277] 10) In the initial operation, there is no semantic error rate resulting from the previous operation. Therefore, in the initial operation, early stop checking is not performed. If there is a previous semantic error for the same data type, the destination is set to a threshold. Determine whether the error rate decreases more than the threshold. If the decrease in the error rate is less than the threshold, the destination performs an early intermediate check operation to determine whether to stop early. If the decrease in the semantic error rate compared to the previously measured semantic error rate is less than the threshold, In this case, the destination determines that the semantic error rate cannot be improved due to retransmission for semantic errors, and accordingly, it may report a system failure, request a new configuration, or request transmission of new data. On the other hand, if the error rate reduction is less than the threshold, If it is greater than that, the destination does not perform an early stop operation and proceeds with the next operation.

[0278] 11) The destination stores the data restored to the semantic level, the attention map obtained based on the prior information transmitted from the source, and the prior information received from the source. Furthermore, the destination transmits to the source the information necessary for semantic error correction. The transmitted information can be broadly categorized into Item A and Item B, as follows.

[0279] A. Attention map generated by the destination to detect semantic errors

[0280] B. A residual area mask, i.e., a residual area attention map, obtained by performing an operation at the destination to extract the remaining area from the attention map transmitted by the source at the current point in time, excluding the area occupied by the attention map generated by the destination.

[0281] The destination determines whether to perform an attention map-related operation on the source using policy A or B, and the attention map-based packet generator (1622) includes attention map information in a packet delivered to the source.

[0282] 12) The source is the number of retransmissions for semantic error correction. Check if it is abnormal. If the current number of retransmissions is In this case, to prevent retransmissions for indiscriminate semantic error correction, the source signals a system failure, initializes the current counter to 0, and flushes the information related to the stored attention map. Afterwards, the source requests a new configuration or requests transmission of new data. Alternatively, the retransmission count counter If less than, the source verifies that a retransmission operation for semantic error correction is possible, increments the counter by one, and proceeds with the following operations.

[0283] 13) The source extracts residual area data required to perform semantic error correction at the destination, depending on the type of attention map received from the destination. For example, if the information received from the destination includes item A, the source generates a residual area attention map, which is a mask for the remaining portion excluding the portion included in the attention map received from the destination, compared to the attention map generated using prior information based on the original data held by the source, and applies the residual area-based data to the original data. As another example, if the source receives item B from the destination, the source generates residual area data by applying the residual area attention map corresponding to the residual area mask to the original data.

[0284] 14) The source performs the same operation as 2) using the residual region-based data as input to the encoder. At this time, the prior information used is prior information for the residual region-based data, and the prior information can be transmitted to the destination.

[0285] In the initial configuration, if semantic error detection is set to be performed on data restored to a semantic level based on the residual region at the destination, a semantic error correction operation is performed when a semantic error occurs in the data, and thus the residual region-based attention map generated at the source is stored. If semantic error detection is not performed on data restored to a semantic level based on the residual region at the destination according to the initial configuration, the source may not perform operations related to generating an attention map based on the data (e.g., data restored to a semantic level based on the residual region). Accordingly, the semantic packet transmitted to the destination may not include the residual region-based attention map, and the residual region-based attention map may not be stored after the semantic packet is transmitted.

[0286] 15) The destination, similar to the process of 5), receives semantic packet(s) and decomposes the packets into a residual local attention map, prior information related to the residual local data, and semantic representation(s). Then, the destination uses the semantic representation(s) as input to the decoder to perform semantic-level restoration, thereby obtaining semantic-level residual local data.

[0287] 16) The destination performs an encoding operation similar to 6) using the restored residual region-based data at the semantic level. Similarly to 6), the network can be trained so that the prior information acquired during the encoding operation can be extracted similarly to the prior information for the residual region-based data transmitted from the source. Thereafter, operations 17) or 18) are performed depending on whether semantic error detection is performed on the residual region-based data.

[0288] 17) When performing semantic error checking on residual region-based data, the destination generates a residual region-based attention map for the residual region-based data restored to the semantic level, similar to the operation of 3), and performs the same operation as 8) using the residual region-based attention map transmitted from the source.

[0289] If the measured semantic error rate is If the average semantic error rate is lower than , the destination determines that the residual region-based data corresponding to the overlapping portion of the attention maps can be used for semantic error correction. Accordingly, the destination synthesizes a portion of the residual region-based data (e.g., the residual region-based data corresponding to the overlapping portion of the residual region-based attention maps obtained from the source and destination) into the existing semantic error correction target data, and performs the operation of 19) below. On the other hand, if the average semantic error rate is In higher cases, the source and destination perform operations 10) to 16) to perform semantic error correction on the remaining region-based data, and then perform operation 17) again.

[0290] 18) If semantic error checking is not performed on the residual region-based data, the destination synthesizes the residual region-based data restored to the semantic level into the semantic error correction target, i.e., the data restored to the semantic level in the previous operation.

[0291] 19) The encoding operation is performed using the composited data from the destination as input. At this time, prior information required for attention map-related operations in the encoder (e.g., prior information generated and transmitted based on the original data) previously received and stored can be used in the learning process to reduce the difference between the two priors.

[0292] Afterwards, the semantic error rate is measured based on the original data-based attention map stored in the destination and the synthesized data attention map obtained through the encoding operation of the destination, for example, as in [Mathematical Equation 9]. If the measured semantic error rate is less than the threshold value, If the error rate is less than 0, the destination can use the semantic representation(s) as input to the task, perform the task operation, and calculate the corresponding loss. A semantic error rate below a certain level indicates that the semantic representation(s) generated based on the synthesized data enable the task located at the destination to operate at the semantic level and perform the desired operation.

[0293] On the other hand, the measured semantic error rate is below the threshold In this case, the destination treats the synthesized data of the current point in time as data restored to the semantic level, thereby designating it as a target for semantic error correction, storing it, and performing semantic error correction using information related to the attention map based on the synthesized data. This can be understood as repeating operation 19) through operations 9) to 16) and operation 17) or 18) described above.

[0294] When inference is performed after learning has been performed, the process can proceed similarly to the process described in learning. The differences between the learning process and the inference process are as follows. First, as described above, the learning process proceeds as follows: the source encoder learns to obtain prior information related to the input data, semantic representation(s), attention map, and prior information transfer; the encoder located at the destination learns to derive values ​​close to the prior information transferred from the source; and attention map generation based on this. In contrast, since the inference process utilizes the network that constitutes the encoder, decoder, and downstream task that have completed learning, the data that transfers semantic information from the source to the destination does not include prior information, and the prior information extracted from each encoder can be used to generate an attention map.

[0295]

[0296] The present disclosure proposes information transfer between a source and a destination and operations between the source and the destination in a system supporting semantic communication, based on an attention map and prior information generated and transferred by the source based on data, and for the destination to perform semantic error detection and correction based on an attention map using prior information. When the source sets a data area that it wants to focus on in relation to the meaning it wants to convey in input data as a semantic element, tasks related to service provision located at the destination can consider the necessary semantic area. Through this, each task derives an attention map using attention-related techniques and data-based prior information to extract the semantic area required for service provision from the input data, thereby enabling the application of a technique to attend to additional areas not covered by existing attention techniques, and accordingly, by simultaneously addressing inductive bias, a more specific semantic area can be expressed in the attention map. In addition, learning is performed to derive prior information close to the priors transmitted from the destination to the source, and by extracting an attention map using the learned priors, it is possible to help derive the attention map close to the data region intended by the source, which can more clearly indicate the attention region for performing semantic error detection and correction, and reduce the number of operations for performing overall semantic error detection and correction.

[0297] By utilizing additional inductive bias through operations at the source and destination using priors, the data set used for training can be reduced. This can reduce the communication overhead associated with forward propagation and back-propagation operations for gradient transfer between the source and destination during training. For example, if the number of input tokens of a model using the attention technique is n, and the number of layers to which priors are applied in an encoder to which the attention technique is applied is ℓ, then in terms of overhead for prior information, which is additional information transmitted from the source to the destination, if distribution information related to the input data is utilized as a prior, n 2 It is possible to reduce the number of ×ℓ priors to 2×n×ℓ. Here, 2 can be interpreted as information related to the center location and window size of the distribution for the input tokens in [Equation 6] and [Equation 7]. This indicates that by additionally transmitting prior information with minimal overhead, the destination can extract an attention map close to the data region for the intended action from the source.

[0298]

[0299] Figure 27 illustrates an example of a wireless device applicable to the present disclosure. The wireless device may be implemented in various forms depending on the use case / service (see Figure 1).

[0300] Referring to FIG. 27, the wireless device (200) corresponds to the wireless device (200) of FIG. 2 and may be composed of various elements, components, units / units, and / or modules. For example, the wireless device (200) may include a communication unit (210), a control unit (220), a memory unit (230), and additional elements (240). The communication unit may include a communication circuit (212) and a transceiver(s) (214). For example, the communication circuit (212) may include one or more processors (202) and / or one or more memories (204) of FIG. 2. For example, the transceiver(s) (214) may include one or more transceivers (206) and / or one or more antennas (208) of FIG. 2. The control unit (220) is electrically connected to the communication unit (210), the memory unit (230), and the additional elements (240) and controls the overall operations of the wireless device. For example, the control unit (220) can control the electrical / mechanical operations of the wireless device based on the program / code / command / information stored in the memory unit (230). In addition, the control unit (220) can transmit information stored in the memory unit (230) to an external device (e.g., another communication device) via a wireless / wired interface through the communication unit (210), or store information received from an external device (e.g., another communication device) via a wireless / wired interface in the memory unit (230).

[0301] The additional element (240) may be configured in various ways depending on the type of the wireless device. For example, the additional element (240) may include at least one of a power unit / battery, an input / output unit (I / O unit), a driving unit, and a computing unit. Although not limited thereto, the wireless device may be implemented in the form of a robot (Fig. 1, 100a), a vehicle (Fig. 1, 100b-1, 100b-2), an XR device (Fig. 1, 100c), a portable device (Fig. 1, 100d), a home appliance (Fig. 1, 100e), an IoT device (Fig. 1, 100f), a digital broadcasting terminal, a hologram device, a public safety device, an MTC device, a medical device, a fintech device (or a financial device), a security device, a climate / environmental device, an AI server / device (Fig. 1, 400), a base station (Fig. 1, 200), a network node, etc. Wireless devices may be mobile or stationary depending on the use / service.

[0302] In FIG. 27, various elements, components, units / parts, and / or modules within the wireless device (200) may be entirely interconnected via a wired interface, or at least some may be wirelessly connected via a communication unit (210). For example, within the wireless device (200), the control unit (220) and the communication unit (210) may be wired, and the control unit (220) and a first unit (e.g., 230, 240) may be wirelessly connected via the communication unit (210). In addition, each element, component, unit / part, and / or module within the wireless device (200) may further include one or more elements. For example, the control unit (220) may be composed of a set of one or more processors. For example, the control unit (220) may be composed of a set of a communication control processor, an application processor, an electronic control unit (ECU), a graphics processing processor, a memory control processor, etc. As another example, the memory unit (130) may be composed of RAM (Random Access Memory), DRAM (Dynamic RAM), ROM (Read Only Memory), flash memory, volatile memory, non-volatile memory, and / or a combination thereof.

[0303] Below, the implementation example of Fig. 27 is described in more detail with reference to the drawings.

[0304] Figure 28 illustrates examples of portable devices applicable to the present disclosure. Portable devices may include smartphones, smart pads, wearable devices (e.g., smartwatches, smartglasses), and portable computers (e.g., laptops, etc.). Portable devices may also be referred to as mobile stations (MS), user terminals (UT), mobile subscriber stations (MSS), subscriber stations (SS), advanced mobile stations (AMS), or wireless terminals (WT).

[0305] Referring to FIG. 28, the portable device (200) may include an antenna unit (208), a communication unit (210), a control unit (220), a memory unit (230), a power supply unit (240a), an interface unit (240b), and an input / output unit (240c). The antenna unit (208) may be configured as a part of the communication unit (210). Blocks 210 to 230 / 240a to 240c of FIG. 28 correspond to blocks 210 to 230 / 240 of FIG. 27, respectively.

[0306] The communication unit (210) can transmit and receive signals (e.g., data, control signals, etc.) with other wireless devices and base stations. The control unit (220) can control components of the portable device (200) to perform various operations. The control unit (220) can include an AP (Application Processor). The memory unit (230) can store data / parameters / programs / codes / commands required for operating the portable device (200). In addition, the memory unit (230) can store input / output data / information, etc. The power supply unit (240a) supplies power to the portable device (200) and can include a wired / wireless charging circuit, a battery, etc. The interface unit (240b) can support connection between the portable device (200) and other external devices. The interface unit (240b) can include various ports (e.g., audio input / output ports, video input / output ports) for connection with external devices. The input / output unit (240c) can input or output video information / signals, audio information / signals, data, and / or information input from a user. The input / output unit (240c) may include a camera, a microphone, a user input unit, a display unit (240d), a speaker, and / or a haptic module.

[0307] For example, in the case of data communication, the input / output unit (240c) obtains information / signals (e.g., touch, text, voice, image, video) input by the user, and the obtained information / signals can be stored in the memory unit (230). The communication unit (210) converts the information / signals stored in the memory into wireless signals, and can directly transmit the converted wireless signals to other wireless devices or to a base station. In addition, the communication unit (210) can receive wireless signals from other wireless devices or base stations, and then restore the received wireless signals to the original information / signals. The restored information / signals can be stored in the memory unit (230) and then output in various forms (e.g., text, voice, image, video, haptic) through the input / output unit (240c).

[0308] Figure 29 illustrates examples of vehicles or autonomous vehicles applicable to the present disclosure. The vehicles or autonomous vehicles may be implemented as mobile robots, cars, trains, manned / unmanned aerial vehicles (AVs), ships, etc.

[0309] Referring to FIG. 29, a vehicle or autonomous vehicle (200-1) may include an antenna unit (208-1), a communication unit (210-1), a control unit (220-1), a driving unit (240a-1), a power supply unit (240b-1), a sensor unit (240c-1), and an autonomous driving unit (240d-1). The antenna unit (208-1) may be configured as a part of the communication unit (210-1). Blocks 210-1 / 230-1 / 240a-1 to 240d-1 of FIG. 29 correspond to blocks 210 / 230 / 240 of FIG. 27, respectively.

[0310] The communication unit (210-1) can transmit and receive signals (e.g., data, control signals, etc.) with external devices such as other vehicles, base stations (e.g., base stations, roadside base stations (ROS), etc.), and servers. The control unit (220-1) can control elements of the vehicle or autonomous vehicle (200-1) to perform various operations. The control unit (220-1) may include an ECU (Electronic Control Unit). The drive unit (240a-1) can drive the vehicle or autonomous vehicle (200-1) on the ground. The drive unit (240a-1) may include an engine, a motor, a power train, wheels, brakes, a steering device, etc. The power supply unit (240b-1) supplies power to the vehicle or autonomous vehicle (200-1) and may include a wired / wireless charging circuit, a battery, etc. The sensor unit (240c-1) can obtain vehicle status, surrounding environment information, user information, etc. The sensor unit (240c-1) may include an IMU (inertial measurement unit) sensor, a collision sensor, a wheel sensor, a speed sensor, an incline sensor, a weight detection sensor, a heading sensor, a position module, a vehicle forward / backward sensor, a battery sensor, a fuel sensor, a tire sensor, a steering sensor, a temperature sensor, a humidity sensor, an ultrasonic sensor, an illuminance sensor, a pedal position sensor, etc. The autonomous driving unit (240d-1) may implement a technology for maintaining a driving lane, a technology for automatically controlling speed such as adaptive cruise control, a technology for automatically driving along a set path, a technology for automatically setting a path and driving when a destination is set, etc.

[0311] For example, the communication unit (210-1) can receive map data, traffic information data, etc. from an external server. The autonomous driving unit (240d-1) can generate an autonomous driving route and driving plan based on the acquired data. The control unit (220-1) can control the drive unit (240a-1) so that the vehicle or autonomous vehicle (200-1) moves along the autonomous driving route according to the driving plan (e.g., speed / direction control). During autonomous driving, the communication unit (210-1) can irregularly / periodically acquire the latest traffic information data from an external server and can acquire surrounding traffic information data from surrounding vehicles. In addition, during autonomous driving, the sensor unit (240c-1) can acquire vehicle status and surrounding environment information. The autonomous driving unit (240d-1) can update the autonomous driving route and driving plan based on newly acquired data / information. The communication unit (210-1) can transmit information regarding the vehicle location, autonomous driving route, driving plan, etc. to an external server. The external server can predict traffic information data in advance using AI technology, etc. based on information collected from the vehicle or autonomous vehicles, and provide the predicted traffic information data to the vehicle or autonomous vehicles. If the device (220-2) is an autonomous vehicle, it can perform the same procedure as the vehicle or autonomous vehicle (200-1). In addition, if the device (220-2) is a base station or a roadside base station, the device (220-2) can transmit data, control signals, etc. to the vehicle or autonomous vehicle (200-1) through the communication unit (210-2).

[0312] Figure 30 illustrates an example of a vehicle applicable to the present disclosure. The vehicle may also be implemented as a means of transportation, a train, an aircraft, a ship, etc. Referring to Figure 30, the vehicle (200) may include a communication unit (210), a control unit (220), a memory unit (230), an input / output unit (240a), and a position measurement unit (240b). Here, blocks 210 to 230 / 240a to 240b correspond to blocks 210 to 230 / 240 of Figure 27, respectively.

[0313] The communication unit (210) can transmit and receive signals (e.g., data, control signals, etc.) with other vehicles or external devices such as base stations. The control unit (220) can control components of the vehicle (200) to perform various operations. The memory unit (230) can store data / parameters / programs / codes / commands that support various functions of the vehicle (100). The input / output unit (240a) can output AR / VR objects based on information in the memory unit (230). The input / output unit (240a) can include a HUD. The position measurement unit (240b) can obtain position information of the vehicle (200). The position information can include absolute position information of the vehicle (200), position information within a driving line, acceleration information, position information with respect to surrounding vehicles, etc. The position measurement unit (240b) can include GPS and various sensors.

[0314] For example, the communication unit (210) of the vehicle (200) can receive map information, traffic information, etc. from an external server and store them in the memory unit (230). The location measurement unit (240b) can obtain vehicle location information through GPS and various sensors and store the information in the memory unit (230). The control unit (220) can create a virtual object based on the map information, traffic information, and vehicle location information, and the input / output unit (240a) can display the created virtual object on the vehicle window (240a-1, 240a-2). In addition, the control unit (220) can determine whether the vehicle (200) is being driven normally within the driving line based on the vehicle location information. If the vehicle (200) abnormally deviates from the driving line, the control unit (220) can display a warning on the vehicle window through the input / output unit (240a). Additionally, the control unit (220) can broadcast a warning message regarding driving abnormalities to surrounding vehicles through the communication unit (210). Depending on the situation, the control unit (220) can transmit vehicle location information and information regarding driving / vehicle abnormalities to relevant authorities through the communication unit (210).

[0315] Figure 31 illustrates examples of XR devices applicable to the present disclosure. The XR devices may be implemented as HMDs, head-up displays (HUDs) installed in vehicles, televisions, smartphones, computers, wearable devices, home appliances, digital signage, vehicles, robots, and the like.

[0316] Referring to FIG. 31, the XR device (200a) may include a communication unit (210), a control unit (220), a memory unit (230), an input / output unit (240a), a sensor unit (240b), and a power supply unit (240c). Here, blocks 210 to 230 / 240a to 240c of FIG. 31 correspond to blocks 210 to 230 / 240 of FIG. 27, respectively.

[0317] The communication unit (210) can transmit and receive signals (e.g., media data, control signals, etc.) with external devices such as other wireless devices, portable devices, or media servers. The media data can include videos, images, sounds, etc. The control unit (220) can control components of the XR device (200a) to perform various operations. For example, the control unit (220) can be configured to control and / or perform procedures such as video / image acquisition, (video / image) encoding, metadata generation and processing, etc. The memory unit (230) can store data / parameters / programs / codes / commands required for driving the XR device (200a) / generating XR objects. The input / output unit (240a) can obtain control information, data, etc. from the outside, and output the generated XR object. The input / output unit (240a) can include a camera, a microphone, a user input unit, a display unit, a speaker, and / or a haptic module. The sensor unit (240b) can obtain the XR device status, surrounding environment information, user information, etc. The sensor unit (240b) may include a proximity sensor, an illuminance sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, and / or a radar. The power supply unit (240c) supplies power to the XR device (200a) and may include a wired / wireless charging circuit, a battery, etc.

[0318] For example, the memory unit (230) of the XR device (200a) may include information (e.g., data, etc.) required for creating an XR object (e.g., AR / VR / MR object). The input / output unit (240a) may obtain a command to operate the XR device (200a) from the user, and the control unit (220) may operate the XR device (200a) according to the user's operating command. For example, when the user attempts to watch a movie, news, etc. through the XR device (200a), the control unit (220) may transmit content request information to another device (e.g., a mobile device (200b)) or a media server through the communication unit (230). The communication unit (230) may download / stream content such as movies and news from another device (e.g., a mobile device (200b)) or a media server to the memory unit (230). The control unit (220) controls and / or performs procedures such as video / image acquisition, (video / image) encoding, and metadata generation / processing for content, and can generate / output an XR object based on information about surrounding space or real objects acquired through the input / output unit (240a) / sensor unit (240b).

[0319] In addition, the XR device (200a) is wirelessly connected to the mobile device (200b) through the communication unit (210), and the operation of the XR device (200a) can be controlled by the mobile device (200b). For example, the mobile device (200b) can act as a controller for the XR device (200a). To this end, the XR device (200a) can obtain 3D location information of the mobile device (200b), and then generate and output an XR object corresponding to the mobile device (200b).

[0320] Figure 32 illustrates examples of robots applicable to the present disclosure. Robots can be classified into industrial, medical, household, and military types, depending on their intended use or field.

[0321] Referring to FIG. 32, the robot (200) may include a communication unit (210), a control unit (220), a memory unit (230), an input / output unit (240a), a sensor unit (240b), and a driving unit (240c). Here, blocks 210 to 230 / 240a to 240c of FIG. 32 correspond to blocks 210 to 230 / 240 of FIG. 27, respectively.

[0322] The communication unit (210) can transmit and receive signals (e.g., driving information, control signals, etc.) with external devices such as other wireless devices, other robots, or control servers. The control unit (220) can control components of the robot (200) to perform various operations. The memory unit (230) can store data / parameters / programs / codes / commands that support various functions of the robot (200). The input / output unit (240a) can obtain information from the outside of the robot (200) and output information to the outside of the robot (200). The input / output unit (240a) can include a camera, a microphone, a user input unit, a display unit, a speaker, and / or a haptic module. The sensor unit (240b) can obtain internal information of the robot (200), surrounding environment information, user information, etc. The sensor unit (240b) may include a proximity sensor, an illuminance sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, a radar, etc. The driving unit (240c) may perform various physical operations such as moving the robot joints. In addition, the driving unit (240c) may enable the robot (200) to drive on the ground or fly in the air. The driving unit (240c) may include an actuator, a motor, wheels, brakes, propellers, etc.

[0323] Figure 33 illustrates an example of an AI device applicable to the present disclosure.

[0324] AI devices can be implemented as fixed or mobile devices, such as TVs, projectors, smartphones, PCs, laptops, digital broadcasting terminals, tablet PCs, wearable devices, set-top boxes (STBs), radios, washing machines, refrigerators, digital signage, robots, and vehicles.

[0325] Referring to FIG. 33, the AI ​​device (200) may include a communication unit (210), a control unit (220), a memory unit (230), an input / output unit (240a / 240b), a learning processor unit (240c), and a sensor unit (240d). Blocks 210 to 230 / 240a to 240d of FIG. 33 correspond to blocks 210 to 230 / 140 of FIG. 27, respectively.

[0326] The communication unit (210) can transmit and receive wired and wireless signals (e.g., sensor information, user input, learning models, control signals, etc.) with external devices such as other AI devices (e.g., 100a to 100f, 120 of FIG. 1) or AI servers (e.g., 100g of FIG. 1) using wired and wireless communication technology. To this end, the communication unit (210) can transmit information within the memory unit (230) to the external device or transfer a signal received from the external device to the memory unit (230).

[0327] The control unit (220) may determine at least one executable operation of the AI ​​device (200) based on information determined or generated using a data analysis algorithm or a machine learning algorithm. In addition, the control unit (220) may control components of the AI ​​device (200) to perform the determined operation. For example, the control unit (220) may request, search, receive, or utilize data from the learning processor unit (240c) or the memory unit (230), and may control components of the AI ​​device (200) to perform at least one executable operation, a predicted operation, or an operation determined to be desirable. In addition, the control unit (220) may collect history information including the operation contents of the AI ​​device (200) or user feedback on the operation, and store the collected history information in the memory unit (230) or the learning processor unit (240c), or transmit the collected history information to an external device such as an AI server (FIG. 1, 100g). The collected history information may be used to update a learning model.

[0328] The memory unit (230) can store data that supports various functions of the AI ​​device (200). For example, the memory unit (230) can store data obtained from the input unit (240a), data obtained from the communication unit (210), output data of the learning processor unit (240c), and data obtained from the sensing unit (140). In addition, the memory unit (230) can store control information and / or software codes necessary for the operation / execution of the control unit (220).

[0329] The input unit (240a) can obtain various types of data from the outside of the AI ​​device (200). For example, the input unit (220) can obtain learning data for model learning, input data to which the learning model will be applied, etc. The input unit (240a) may include a camera, a microphone, and / or a user input unit. The output unit (240b) may generate output related to vision, hearing, or touch. The output unit (240b) may include a display unit, a speaker, and / or a haptic module, etc. The sensing unit (140d) can obtain at least one of internal information of the AI ​​device (200), information about the surrounding environment of the AI ​​device (200), and user information using various sensors. The sensing unit (140d) may include a proximity sensor, an illuminance sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, and / or a radar, etc.

[0330] The learning processor unit (240c) can train a model composed of an artificial neural network using learning data. The learning processor unit (240c) can perform AI processing together with the learning processor unit of the AI ​​server (Fig. 1, 100g). The learning processor unit (240c) can process information received from an external device via the communication unit (210) and / or information stored in the memory unit (230). In addition, the output value of the learning processor unit (240c) can be transmitted to an external device via the communication unit (210) and / or stored in the memory unit (230).

[0331]

[0332] It is clear that the examples of the proposed methods described above can also be considered as a type of proposed methods, as they can be included as one of the implementation methods of the present disclosure. Furthermore, the proposed methods described above can be implemented independently, but they can also be implemented in the form of a combination (or merge) of some of the proposed methods. Information regarding the applicability of the proposed methods (or information regarding the rules of the proposed methods) can be defined by a rule such that the base station notifies the terminal of the application of the proposed methods through a predefined signal (e.g., a physical layer signal or a higher layer signal).

[0333] The present disclosure may be embodied in other specific forms without departing from the technical ideas and essential features described herein. Therefore, the above detailed description should not be construed as limiting in all respects but rather as illustrative. The scope of the present disclosure should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present disclosure are intended to be included within the scope of the present disclosure. Furthermore, claims that do not explicitly cite each other in the claims may be combined to form embodiments or incorporated into new claims through post-filing amendments.

[0334] Embodiments of the present disclosure can be applied to various wireless access systems. As an example of various wireless access systems, 3GPP (3 rd Generation Partnership Project) or 3GPP2 system, etc.

[0335] The embodiments of the present disclosure can be applied not only to the various wireless access systems described above, but also to all technical fields that utilize these various wireless access systems. Furthermore, the proposed method can also be applied to mmWave and THz communication systems utilizing ultra-high frequency bands.

[0336] Additionally, embodiments of the present disclosure can be applied to various applications such as autonomous vehicles and drones.

Claims

1. A method performed by a first device in a wireless communication system, Step of establishing a connection with a second device; A step of receiving a first message requesting capability information from the second device; A step of transmitting a second message including the above capability information to the second device; A step of receiving setting information for communication from the second device; and A step of learning for semantic communication supporting at least one task based on the above setting information or performing the semantic communication, A method in which the above setting information includes information for performing semantic error detection and correction using an attention map generated based on prior information of input data.

2. In claim 1, A method in which the above-mentioned configuration information includes at least one of information regarding whether fine-tuning is performed, information regarding a model of an encoder and decoder for the semantic communication, information regarding a technique for generating the prior information, information regarding a technique for transmitting the prior information, information regarding whether a bit-wise attention map is set, information regarding a threshold related to the bit-wise attention map, information regarding a threshold related to error detection, information regarding a maximum semantic data retransmission count, information regarding an early termination threshold, and information regarding whether semantic error detection is performed for residual region-based data.

3. In claim 1, A method in which the attention map generated based on the above prior information is generated by fusion of an initial attention map based on self-attention and the priors related to structural information about the input data.

4. In claim 3, The above fusion is a method including an operation of weighting and adding an attention matrix generated by a first layer and an attention logit matrix generated by a second layer, which is a layer preceding the first layer, among layers that perform operations related to an attention map included in a network for generating the attention map.

5. In claim 3, The above fusion is a method including an operation of performing the softmax operation after adding information related to the distribution of the input data and information related to the relative position of each token to the target of the softmax operation for the self-attention.

6. In claim 1, The step of performing the above learning or the above semantic communication is, A step of transmitting a first semantic packet for a task using the above input data; A step of receiving feedback information for semantic error correction, which is related to a residual area which is part of the above input data; and A method comprising the step of transmitting a second semantic packet related to the remaining area.

7. In claim 6, The step of transmitting the first semantic packet is: A step of generating a first semantic representation for the task from the input data; and A method comprising the step of transmitting, to the second device, the first semantic packet including the first attention map corresponding to the first semantic expression, the first semantic expression, and the prior.

8. In claim 6, A method wherein the feedback information comprises a second attention map generated for semantic error detection in the second device, or a residual region attention map indicating a remaining area excluding a portion occupied by the second attention map in the first attention map.

9. In claim 6, The step of transmitting the second semantic packet is: generating a second semantic representation for the task from residual region data, which is part of the input data indicated by the residual region; and A method comprising the step of transmitting a second semantic packet including second attention information corresponding to the second semantic expression, the second semantic expression, and a prior of the residual region data to the second device.

10. In a method performed by a second device in a wireless communication system, Step of establishing a connection with the first device; A step of transmitting a first message requesting capability information to the first device; A step of receiving a second message including the above capability information from the first device; A step of transmitting setting information for communication to the first device; and A step of performing semantic communication that supports at least one task based on the above setting information, A method in which the above setting information includes information for performing semantic error detection and correction using an attention map generated based on prior information of input data.

11. In claim 10, The steps for performing the above semantic communication are: A step of receiving a semantic representation from the first device and a first prior generated by the first device; A step of obtaining data restored to a semantic level by decoding the above semantic expression; A method comprising a step of performing learning to reduce the difference between the second prior obtained through encoding in the second device and the first prior using data restored to the semantic level.

12. In claim 11, A method wherein the difference between the second prior and the first prior comprises a difference between a first attention logit matrix received from the first device and a second attention logit matrix generated from the second device.

13. In claim 11, A method wherein the difference between the second fryer and the first fryer comprises a difference between a first center position for a distribution of fryer information generated from the first device and a first window size set based on the first center position and a second center position for a distribution of fryer information generated from the second device and a second window size set based on the second center position.

14. In claim 10, The steps for performing the above semantic communication are: A step of receiving a first packet including first semantic information for at least one task; A step of obtaining data restored to a first semantic level by performing semantic decoding on a first semantic expression included in the first packet; A step of obtaining a second prior-based attention map by performing semantic encoding on data restored to the first semantic level; A step of transmitting feedback information for semantic error correction related to a residual area, which is part of the input data, based on the attention map based on the second prior; A step of receiving a second packet including second semantic information related to the remaining area; and A method comprising a step of performing semantic error correction based on the second semantic information included in the second packet.

15. In claim 14, The step of transmitting the above feedback information is: If the difference between the first prior-based attention map and the second prior-based attention map included in the first packet is greater than a threshold, a step of checking the remaining region corresponding to the difference; and A method comprising the step of transmitting the feedback information indicating at least one of a mask indicating the residual region or an attention map based on the second prior.

16. In claim 15, A method in which the residual region includes the remainder of the first prior-based attention map generated by the first device, excluding the overlapping region of the first prior-based attention map and the second prior-based attention map.

17. In claim 11, The steps for performing the above semantic error correction are: A step of obtaining data restored to a second semantic level for the residual region by performing semantic decoding on a second semantic expression included in the second packet; A method comprising a step of synthesizing data restored to the first semantic level and data restored to the second semantic level.

18. In a first device in a wireless communication system, Transmitter and receiver; and A processor connected to the above transmitter and receiver is included, The above processor, Establish a connection with the second device, Receive a first message requesting capability information from the second device, Transmitting a second message including the above capability information to the second device, Receive setup information for communication from the second device, Controls semantic communication to support multiple tasks based on the above setting information, The above setting information is a first device including information for performing semantic error detection and correction using an attention map generated based on prior information of input data.

19. In communication devices, At least one processor; At least one computer memory connected to said at least one processor and storing instructions that direct operations when executed by said at least one processor, The above actions are, A step of establishing a connection with another communication device; A step of receiving a first message requesting capability information from said other communication device; A step of transmitting a second message including the above capability information to the other communication device; A step of receiving setting information for communication from the other communication device; and A step of performing semantic communication that supports multiple tasks based on the above setting information is included. A method in which the above setting information includes information for performing semantic error detection and correction using an attention map generated based on prior information of input data.

20. In a non-transitory computer-readable medium storing at least one instruction, comprising at least one instruction executable by the processor, At least one of the above commands causes the device to: Establish a connection with another device, Receive a first message requesting capability information from said other device, Transmitting a second message including the above capability information to the other device, Receive setup information for communication from the other device, Controls semantic communication to support multiple tasks based on the above setting information, The above setting information is a non-transitory computer-readable medium including information for performing semantic error detection and correction using an attention map generated based on prior information of input data.

Citation Information

Patent Citations

  • Recurrent networks with motion-based attention for video understanding

    KR1020180111959A

  • Apparatus and method for generating transmit and receive signals in wireless communication system

    WO2024034695A1

  • Device and method for detecting semantic error in wireless communication system

    WO2024162500A1

Cited By

  • Cloud-based real-time messaging layer for resource transmission

    US12665699B1