Resource occupation method and communication apparatus
By collaboratively determining the memory resource usage status of AI model data, the problem of limited computing power of network devices was solved, and computing efficiency and memory resource utilization were improved.
Patent Information
- Application Number
- PCT/CN2025/109887
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-14
- Filing Date
- 2025-07-22
- Publication Date
- 2026-02-19
AI Technical Summary
Network devices have limited computing power, and the frequent loading of model data from storage devices increases the runtime of computing tasks. How can we improve the computing efficiency of network devices?
By coordinating with terminal and network devices, the memory resource usage status of AI model data is determined, and either a continuous or temporary usage status is adopted. Other data can be allowed or prohibited from occupying memory resources, thereby reducing the time to load data from storage devices.
It improves the computing efficiency of network devices, optimizes the utilization of memory resources, and reduces the overall running time of computing tasks.
Smart Images

Figure CN2025109887_19022026_PF_FP_ABST
Abstract
Description
Method and communication apparatus for resource occupation
[0001] The present application claims priority to the Chinese patent application No. 202411120796.4, filed on August 14, 2024, and entitled "Method and communication apparatus for resource occupation", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of communication, and more particularly, to a method and communication apparatus for resource occupation. BACKGROUND
[0003] With the continuous development and improvement of extended reality (XR) technology, the accuracy requirements for object detection and recognition of image content in a picture are also increasing. In order to improve the accuracy of object detection and recognition, deep neural network (DNN) is introduced. However, with the development of communication, the computing tasks required by DNN are becoming more and more heavy.
[0004] Currently, terminal devices can share a part of the computing tasks of DNN, and with the development of large language models, network devices can also perform the computing tasks of DNN. However, due to the limitations of network device deployment, the computing capability of network devices is limited, and network devices need to frequently load model data from the storage devices of network devices when performing computing, which will cause the running time of the computing tasks of network devices to increase.
[0005] Therefore, how to improve the computing efficiency of network devices is a problem to be solved. SUMMARY
[0006] The present application provides a communication method and a communication apparatus to reduce the time for network devices to load model data from storage devices and improve the computing efficiency of network devices.
[0007] In a first aspect, a communication method is provided, which can be applied to a terminal side, such as a terminal or a communication module / processing module in the terminal, or a circuit or chip responsible for communication function in the terminal (such as a modem chip, also known as a baseband chip, or a system on chip (SoC) chip or system in package (SIP) chip containing a modem core), or a circuit or chip responsible for processing function in the terminal (such as a graphics processing unit (GPU)). The following describes the method applied to a terminal as an example.
[0008] In the method, the terminal sends first request information, the first request information being used for requesting an artificial intelligence (AI) model data to occupy a memory resource; the terminal receives first indication information, and determines, based on the first indication information, a memory resource occupation state of the AI model data, the memory resource occupation state being one of candidate occupation states, the candidate occupation states including a persistent occupation state and a temporary occupation state, in the persistent occupation state, the memory resource occupied by the AI model data is prohibited to be occupied by data other than the AI model data, and in the temporary occupation state, the memory resource occupied by the AI model data is allowed to be occupied by data other than the AI model data.
[0009] By using the method, the terminal initiates a request for the AI model data to occupy the memory resource, and receives the indication information, based on which the terminal can determine the memory resource occupation state of the AI model data, which can be one of the persistent occupation state and the temporary occupation state, so that for the memory resource occupied by the AI model data in the persistent occupation state, other data is prohibited to occupy the memory resource, thereby reducing the time for loading data from a storage device, and improving the computing efficiency of a network device; and for the memory resource occupied by the AI model data in the temporary occupation state, other data is allowed to occupy the memory resource, thereby reducing the waste of resources.
[0010] In a possible design, in the persistent occupation state, a time length for which the AI model data occupies the memory resource is within a first time length, and the memory resource is prohibited to be occupied by data other than the AI model data.
[0011] In a possible design, in the persistent occupation state, a time length for which the AI model data occupies the memory resource is less than the first time length, and the memory resource is prohibited to be occupied by data other than the AI model data.
[0012] By using the method, in the persistent occupation state, the AI model data occupies the memory resource within the first time length, and after the first time length, the memory resource can be occupied by data other than the AI model data, so that the AI model data does not always occupy the memory resource, and when the terminal needs, data other than the AI model data can also occupy the memory resource, thereby improving the utilization rate of the memory resource.
[0013] In a possible design, in the temporary occupation state, after a time length for which the AI model data occupies the memory resource is greater than or equal to a second time length, the memory resource is allowed to be occupied by data other than the AI model data.
[0014] In a possible design, in the temporary occupation state, the memory resource is allowed to be occupied by data other than the AI model data after the AI model data occupies the memory resource for a time period longer than a second time period.
[0015] By using the method, when data of a computing task with a higher priority than the AI model computing task arrives in the temporary occupation state, the memory resource can be occupied by the data of the computing task with the higher priority, and the utilization of the memory resource is improved.
[0016] In a possible design, the first request information is specifically used to request any of the following: the AI model data occupies the memory resource in the persistent occupation state; the AI model data occupies the memory resource in the temporary occupation state; the AI model data occupies the memory resource in the persistent occupation state and a time period for occupying the memory resource in the persistent occupation state; and the AI model data occupies the memory resource in the temporary occupation state and a time period for occupying the memory resource in the temporary occupation state.
[0017] In a possible design, determining the memory resource occupation state of the AI model data based on the first indication information comprises: determining, based on the first indication information, that the memory resource occupation state of the AI model data is the persistent occupation state; or determining, based on the first indication information, that the memory resource occupation state of the AI model data is the persistent occupation state and a time period for occupying the memory resource in the persistent occupation state; or determining, based on the first indication information, that the memory resource occupation state of the AI model data is the temporary occupation state; or determining, based on the first indication information, that the memory resource occupation state of the AI model data is the temporary occupation state and a time period for occupying the memory resource in the temporary occupation state.
[0018] By using the above method, the network device can comprehensively consider its own conditions and the priority of the computing task, and send the first indication information to the terminal to inform the terminal of the memory resource occupation state that the network device can allow, so that the computing task with a higher priority can occupy the memory resource, the time for loading data from the storage device is reduced, the overall running time of the computing task is reduced, and the computing efficiency of the network device is improved.
[0019] In a possible design, the candidate occupation states further include a prohibited occupation state, in which the AI model data is prohibited from occupying the memory resource.
[0020] In a possible design, determining the memory resource occupation state of the AI model data based on the first indication information comprises: determining, based on the first indication information, that the memory resource occupation state of the AI model data is the prohibited occupation state.
[0021] In a possible design, the method further includes: determining a first time interval based on the first indication information; and sending second request information after the first time interval, where the second request information is used to request the AI model data to occupy the memory resource.
[0022] In the above manner, the network device can comprehensively consider its own conditions, and when the conditions are insufficient, the network device sends the first indication information to the terminal to inform the terminal of the memory resource occupation state that is prohibited by the network device, and to inform the terminal of the time interval for initiating the request again, so that the terminal initiates the request at intervals and reduces power consumption.
[0023] In a possible design, when the terminal is in a radio resource control (RRC) connected state, the first request information is specifically used to request the AI model data to occupy the memory resource in the persistent occupation state.
[0024] In a possible design, when the terminal is in an RRC connected state or an RRC inactive state, the first request information is specifically used to request the AI model data to occupy the memory resource in the temporary occupation state.
[0025] In the above manner, by associating the RRC state of the terminal with the occupation state of the memory resource, when the RRC state of the terminal changes, the occupation state of the memory resource also changes, thereby improving the utilization rate of the memory resource.
[0026] In a second aspect, a communication method is provided, which is applied to a network side, for example, an access network device of the network side, a module (for example, a circuit, a chip, or a chip system, etc.) in the access network device, or a logic node, a logic module, or software capable of realizing all or part of the functions of the access network device. Hereinafter, the method is taken as an example for description, and the network device can be an access network device or a base station.
[0027] In the method, the network device receives first request information, where the first request information is used to request artificial intelligence (AI) model data to occupy a memory resource; and the network device sends first indication information, where the first indication information is used to indicate an occupation state of the memory resource of the AI model data, and the occupation state of the memory resource is one of candidate occupation states, where the candidate occupation states include a persistent occupation state and a temporary occupation state, in the persistent occupation state, the memory resource occupied by the AI model data is prohibited to be occupied by data other than the AI model, and in the temporary occupation state, the memory resource occupied by the AI model data is allowed to be occupied by data other than the AI model data.
[0028] In a possible design, in the persistent occupation state, the AI model data occupies the memory resource for a time period within a first time period, and the memory resource is prohibited from being occupied by data other than the AI model data.
[0029] In a possible design, in the persistent occupation state, the AI model data occupies the memory resource for a time period less than the first time period, and the memory resource is prohibited from being occupied by data other than the AI model data.
[0030] In a possible design, in the temporary occupation state, after the AI model data occupies the memory resource for a time period greater than or equal to a second time period, the memory resource is allowed to be occupied by data other than the AI model data.
[0031] In a possible design, in the temporary occupation state, after the AI model data occupies the memory resource for a time period greater than the second time period, the memory resource is allowed to be occupied by data other than the AI model data.
[0032] In a possible design, the first request information is specifically used to request any one of the following:
[0033] The AI model data occupies the memory resource in the persistent occupation state;
[0034] The AI model data occupies the memory resource in the temporary occupation state;
[0035] The AI model data occupies the memory resource in the persistent occupation state, and a time period for occupying the memory resource in the persistent occupation state;
[0036] The AI model data occupies the memory resource in the temporary occupation state, and a time period for occupying the memory resource in the temporary occupation state.
[0037] In a possible design, the first indication information is used to indicate the memory resource occupation state of the AI model data, and includes:
[0038] The first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the persistent occupation state; or
[0039] The first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the persistent occupation state, and is specifically used to indicate a time period for occupying the memory resource in the persistent occupation state; or
[0040] The first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the temporary occupation state; or
[0041] The first indication information is specifically used for indicating that the memory resource occupation state of the AI model data is the temporary occupation state, and is specifically used for indicating a time at which the AI model data occupies the memory resource in the temporary occupation state.
[0042] In a possible design, the candidate occupation state further includes a forbidden occupation state, in which the AI model data is forbidden to occupy the memory resource.
[0043] In a possible design, the first indication information is used for indicating the memory resource occupation state of the AI model data, and includes that the first indication information is specifically used for indicating that the memory resource occupation state of the AI model data is the forbidden occupation state.
[0044] In a possible design, the first indication information is further used for indicating a first time interval, and the method further includes: receiving second request information after the first time interval, the second request information being used for requesting the AI model data to occupy the memory resource.
[0045] In a possible design, when the terminal is in a radio resource control (RRC) connected state, the first request information is specifically used for requesting the AI model data to occupy the memory resource in the continuous occupation state.
[0046] In a possible design, when the terminal is in an RRC connected state or an RRC inactive state, the first request information is specifically used for requesting the AI model data to occupy the memory resource in the temporary occupation state.
[0047] The beneficial effects of the above-mentioned second aspect and some implementation manners can be referred to the description of the first aspect and some implementation manners, which will not be described here.
[0048] In a third aspect, a communication apparatus is provided, which has the function of implementing the first aspect, for example, the communication apparatus includes a module or unit or means corresponding to the operations of the first aspect, which can be implemented by software, or by hardware, or by a combination of software and hardware.
[0049] In a possible design, the communication apparatus includes: a transceiver, configured to send first request information, where the first request information is used to request an artificial intelligence (AI) model data to occupy memory resources; and the transceiver is further configured to receive first indication information; and a processing unit, configured to determine, based on the first indication information, a memory resource occupation state of the AI model data, where the memory resource occupation state is one of candidate occupation states, and the candidate occupation states include a persistent occupation state and a temporary occupation state, in the persistent occupation state, the memory resources occupied by the AI model data are prohibited to be occupied by data other than the AI model data, and in the temporary occupation state, the memory resources occupied by the AI model data are allowed to be occupied by data other than the AI model data.
[0050] The transceiver can perform the receiving and sending in the first aspect, and the processing unit can perform other processing in the first aspect other than the receiving and sending.
[0051] The communication apparatus can be a terminal, a communication module in the terminal, or a chip responsible for a communication function, such as a modem chip (also referred to as a baseband chip) or a SoC or SIP chip including a modem module.
[0052] In a fourth aspect, a communication apparatus is provided. The communication apparatus has the functions of the second aspect, for example, the communication apparatus includes modules or units or means corresponding to the operations of the second aspect, which can be implemented in software, or implemented in hardware, or implemented in a combination of software and hardware.
[0053] In a possible design, the communication apparatus includes: a transceiver, configured to receive first request information, where the first request information is used to request an artificial intelligence (AI) model data to occupy memory resources; and the transceiver is further configured to send first indication information, where the first indication information is used to indicate a memory resource occupation state of the AI model data, and the memory resource occupation state is one of candidate occupation states, and the candidate occupation states include a persistent occupation state and a temporary occupation state, in the persistent occupation state, the memory resources occupied by the AI model data are prohibited to be occupied by data other than the AI model data, and in the temporary occupation state, the memory resources occupied by the AI model data are allowed to be occupied by data other than the AI model data.
[0054] The transceiver can perform the receiving and sending in the second aspect, and the communication apparatus further includes a processing unit, which can perform other processing in the second aspect other than the receiving and sending.
[0055] The communication apparatus can be an access network device, or a module (e.g., a circuit, a chip or a chip system, etc.) in the access network device, or a logic node, a logic module or software capable of realizing all or part of the functions of the access network device.
[0056] In a fifth aspect, a communication apparatus is provided. The communication apparatus includes an interface circuit and one or more processors. The one or more processors are coupled to a memory. The memory is configured to store part or all of the computer program or instructions necessary to implement the functions related to any of the first aspect to the second aspect. The one or more processors are configured to execute the computer program or instructions, which when executed cause the communication apparatus to implement the method in any possible design or implementation of the first aspect to the second aspect. The interface circuit is configured to implement the communication function within the communication apparatus and / or the communication function between the communication apparatus and other apparatuses or components.
[0057] In a possible design, the processor is configured to communicate with other apparatuses or components via the interface circuit.
[0058] In a possible design, the communication apparatus can further include the memory.
[0059] The communication apparatus can be a terminal, or a communication module in the terminal, or a chip responsible for the communication function in the terminal, such as a modem chip (also referred to as a baseband chip) or a SoC or SIP chip containing a modem module.
[0060] The communication apparatus can be an access network device, or a module (e.g., a circuit, a chip or a chip system, etc.) in the access network device, or a logic node, a logic module or software capable of realizing all or part of the functions of the access network device.
[0061] In a sixth aspect, a communication system is provided. The communication system includes at least one of the communication apparatuses in the third aspect to the fourth aspect.
[0062] In a seventh aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer program codes or instructions, which when executed by a computer, cause the method in any possible implementation of the first aspect to the second aspect to be implemented.
[0063] In an eighth aspect, a computer program product is provided. The computer program product includes computer program codes or instructions, which when executed by a computer, cause the method in any possible implementation of the first aspect to the second aspect to be implemented.
[0064] In a ninth aspect, a computer program is provided. When the computer program is run, the method in any possible implementation of the first aspect to the second aspect is implemented.
[0065] It should be understood that the beneficial effects of the third aspect to the ninth aspect described above can refer to the first aspect to the second aspect and any possible implementation thereof, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS
[0066] FIG. 1 is a schematic diagram of a communication system suitable for embodiments of the present application.
[0067] FIG. 2 is a schematic diagram of AI model deployment according to embodiments of the present application.
[0068] FIG. 3 is a schematic diagram of DNN model calculation division according to embodiments of the present application.
[0069] FIG. 4 is a schematic diagram of network device calculation according to embodiments of the present application.
[0070] FIG. 5 is a schematic diagram of a communication method according to embodiments of the present application.
[0071] FIG. 6 is a schematic diagram of a transition between RRC states according to embodiments of the present application.
[0072] FIG. 7 is a possible exemplary block diagram of a communication apparatus according to embodiments of the present application.
[0073] FIG. 8 is a schematic diagram of a structure of a terminal according to embodiments of the present application. DETAILED DESCRIPTION
[0074] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0075] Before introducing the solutions of the present application, the following points are explained.
[0076] (1) In the present application, “indication” can include direct indication, indirect indication, explicit indication, implicit indication, and the like. When describing that certain indication information indicates A, it can be understood that the indication information carries A, carries an identifier of A, carries B having an association relationship with A, carries an identifier of B having an association relationship with A, and the like. In other words, if the receiving side of certain indication information can determine A according to the indication information, it can be described that the indication information indicates A, and the specific determination is not limited. When it is understood that the indication information carries A, “indication” can be replaced by “includes”, at this time, similar to the expression “sending / receiving indication information, the indication information indicating A”, it can be replaced by “sending / receiving A”.
[0077] In the present application, the information indicated by the indication information is referred to as to-be-indicated information. In the specific implementation process, there are many ways to indicate the to-be-indicated information, for example, but not limited to, the to-be-indicated information can be directly indicated, such as the to-be-indicated information itself or the index of the to-be-indicated information. The to-be-indicated information can also be indirectly indicated by indicating other information, wherein the other information and the to-be-indicated information have an association relationship. The to-be-indicated information can also be only indicated in part, and the other part of the to-be-indicated information is known or agreed in advance. For example, the indication of a specific information can also be achieved by means of the arrangement order of each information agreed in advance (for example, the protocol stipulates), thereby reducing the indication overhead to a certain extent. In addition, the to-be-indicated information can be sent as a whole, or can be sent separately into multiple sub-information, and the sending period and / or sending time of these sub-information can be the same or different.
[0078] (2) In the present application, the expression " / " is used to represent that the objects associated before and after are in an "or" relationship; for example, A / B can represent A or B. The expression "and / or" is used to represent that the objects associated before and after can be in an and relationship or an or relationship; for example, A and / or B can represent the following cases: A exists alone, B exists alone, A and B exist together, wherein A and B can be single or multiple. "At least one of the following" or similar expressions are used to represent any combination of the listed items; for example, at least one of A, B and (or) C can represent the following cases: A exists alone, B exists alone, C exists alone, A and B exist together, B and C exist together, A and C exist together, A, B and C exist together, wherein A, B and C can be single or multiple.
[0079] (3) In the present application, "sending" and "receiving" represent the direction of signal transmission. For example, "sending information to XX" can be understood as that the destination of the information is XX, which can include direct sending through the air interface, or indirect sending through the air interface by other units or modules. "Receiving information from YY" can be understood as that the source of the information is YY, which can include direct receiving from YY through the air interface, or indirect receiving from YY through the air interface from other units or modules. "Sending" can also be understood as "output" of the chip interface, and "receiving" can also be understood as "input" of the chip interface. In other words, sending and receiving can be carried out between devices, for example, between network devices and terminal devices, or can be carried out within a device, for example, between components, modules, chips, software modules or hardware modules within a device through a bus, wire or interface.
[0080] (4) In the various embodiments of this application, unless otherwise specified or in case of logical conflict, the terms and / or descriptions of different embodiments are consistent and can be referenced by each other. The technical features of different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0081] (5) In this application, "first," "second," and "#1," "#2," and "#A" are merely for descriptive convenience and are used to distinguish objects, and are not intended to limit the scope of the embodiments of this application. They are not used to describe the order or sequence of features. It should be understood that such described objects can be interchanged where appropriate in order to describe solutions other than those in the embodiments of this application.
[0082] (6) In this application, "predefined" can mean a standard protocol predefined, or it can mean a pre-agreed or pre-negotiated agreement between devices. Here, "protocol" can refer to a standard protocol in the field of communications, for example, it may include fourth-generation (4G) protocols. th Generation 4G network, fifth generation (5G) network th This application does not limit the scope to network protocols such as generation (5G), NR, 5.5G, future communication network protocols, and related protocols applied in future communication systems.
[0083] (7) In this application, the words “exemplary,” “for example,” etc., are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as an “example” in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word “example” is intended to present the concept in a concrete manner. In the embodiments of this application, “of,” “corresponding, relevant,” and “corresponding” may sometimes be used interchangeably, and it should be noted that their intended meanings are consistent unless their distinction is emphasized.
[0084] First, let me introduce the communication system to which this application applies.
[0085] The technical solutions provided in the application can be applied to various communication systems, for example, a long term evolution (LTE) system, an LTE frequency division duplex (FDD) system, an LTE time division duplex (TDD), a 5th generation (5G) system or a new radio (NR), and future communication systems, vehicle-to-X (V2X), which can include vehicle-to-network (V2N), vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), vehicle-to-pedestrian (V2P), inter-vehicle communication long term evolution (LTE-V), Internet of Vehicles, machine type communication (MTC), Internet of Things (IoT), inter-machine communication long term evolution (LTE-M), machine-to-machine (M2M), inter-satellite communication, and non-terrestrial network (NTN) systems, and the like.
[0086] As an example, a satellite communication system includes a satellite base station and a terminal device. The satellite base station provides communication services for the terminal device. The satellite base station can also communicate with a base station. The satellite can act as a base station and also as a terminal device. The satellite can refer to a drone, a hot air balloon, a low earth orbit satellite, a medium earth orbit satellite, a high earth orbit satellite, and the like. The satellite can also refer to a non-ground base station or a non-ground device, and the like.
[0087] As an example, V2X communication can include vehicle-to-vehicle (V2V) communication, vehicle-to-infrastructure (V2I) communication, vehicle-to-pedestrian (V2P) communication, and vehicle-to-network (V2N) communication.
[0088] A device in a communication system can send a signal to or receive a signal from another device. Wherein the signal can include information, signaling or data, etc. Wherein the device can also be replaced by an entity, a network entity, a communication device, a communication module, a node, a communication node, etc. The device is taken as an example for description in the embodiments of the present application.
[0089] Referring to FIG. 1, as an example, FIG. 1 is a schematic diagram of a communication system suitable for embodiments of the present application. As shown in FIG. 1, the communication system 10 includes a radio access network (RAN) 100 and a core network (CN) 200. The RAN 100 includes at least one RAN node (e.g., 110a and 110b in FIG. 1, collectively referred to as 110) and at least one terminal (e.g., 120a-120j in FIG. 1, collectively referred to as 120). The RAN 100 can also include other RAN nodes, such as wireless relay devices and / or wireless backhaul devices (not shown in FIG. 1), etc. The terminal 120 is connected to the RAN node 110 in a wireless manner. The RAN node 110 is connected to the CN 200 in a wireless or wired manner. The core network device in the CN 200 and the RAN node 110 in the RAN 100 can be different physical devices respectively, or can be the same physical device integrated with the logical functions of the core network and the logical functions of the radio access network.
[0090] The RAN 100 can be a 3rd generation partnership project (3GPP) related cellular system, such as a fourth generation (4G) mobile communication system, a fifth generation (5G) mobile communication system, or a future communication network. The RAN 100 can also be an open radio access network (O-RAN or ORAN), a cloud radio access network (CRAN), or a wireless fidelity (WiFi) system. The RAN 100 can also be a communication system in which two or more of the above systems are fused.
[0091] The RAN node 110, which can also be referred to as an access network device, a RAN entity, or an access node, etc., forms part of the communication system, and is configured to facilitate the wireless access by the terminals. The RAN nodes 110 in the communication system can be nodes of the same type or nodes of different types. In some scenarios, the roles of the RAN node 110 and the terminal 120 are relative, e.g., the network element 120i in Figure 1 can be a helicopter or a drone, which can be configured to be a mobile base station, and for a terminal 120j accessing the RAN 100 via the network element 120i, the network element 120i is a base station; but for the base station 110a, the network element 120i is a terminal. Both the RAN node 110 and the terminal 120 are sometimes referred to as communication apparatuses, e.g., the network elements 110a and 110b in Figure 1 can be understood as communication apparatuses with base station functionalities, and the network elements 120a-120j can be understood as communication apparatuses with terminal functionalities.
[0092] In a possible scenario, the RAN node can be a base station, an evolved Node B (eNodeB), an access point (AP), a transmission reception point (TRP), a next generation Node B (gNB), a next generation base station in a future mobile communication system, a base station in a future mobile communication system, or an access node in a WiFi system, etc. The RAN node can be a macro base station (e.g., 110a in Figure 1), a micro base station or an indoor station (e.g., 110b in Figure 1), a relay node or a donor node, or a wireless controller in a CRAN scenario. Optionally, the RAN node can also be a server, a wearable device, a vehicle or a vehicle-mounted device, etc. For example, the access network device in vehicle to everything (V2X) technology can be a road side unit (RSU). All or part of the functions of the RAN node in the present application can also be implemented by software functions running on hardware, or by virtualized functions instantiated on a platform (e.g., a cloud platform). The RAN node can also be provided with a communication module, circuit or chip for performing corresponding communication functions, and program instructions for performing corresponding communication functions. The RAN node in the present application can also be a logic node, a logic module or software capable of implementing all or part of the functions of the RAN node.
[0093] In another possible scenario, a terminal is assisted by multiple RAN nodes to implement wireless access, and different RAN nodes respectively implement part of functions of a base station. For example, a RAN node can be a central unit (CU), a distributed unit (DU), a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU), etc. The CU and the DU can be separately arranged, or can also be included in the same network element, for example, in a baseband unit (BBU). The RU can be included in a radio frequency device or a radio frequency unit, for example, included in a remote radio unit (RRU), an active antenna processing unit (AAU), or a remote radio head (RRH).
[0094] In different systems, the CU (or CU-CP and CU-UP), DU or RU can also have different names, but those skilled in the art can understand their meanings. For example, in an ORAN system, the CU can also be referred to as an O-CU (open CU), the DU can also be referred to as an O-DU, the CU-CP can also be referred to as an O-CU-CP, the CU-UP can also be referred to as an O-CU-UP, and the RU can also be referred to as an O-RU. For the convenience of description, the CU, CU-CP, CU-UP, DU and RU are taken as examples for description in this application. Any one of the CU (or CU-CP, CU-UP), DU and RU in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.
[0095] A terminal 120 can be a device or module with corresponding communication functions and capable of accessing the above communication system. The terminal can also be referred to as a terminal device, user equipment (UE), mobile station, mobile terminal, etc. The terminal can be widely applied to various scenarios, such as device-to-device (D2D) communication, vehicle-to-everything (V2X) communication, machine-type communication (MTC), internet of things (IOT), virtual reality, augmented reality, industrial control, automatic driving, remote medical treatment, smart power grid, smart home, smart office, smart wear, smart transportation, smart city, etc. The terminal can be a mobile phone, tablet computer, computer with wireless transceiver function, wearable device, vehicle, unmanned aerial vehicle, helicopter, airplane, ship, robot, mechanical arm, smart home device, transport vehicle with wireless communication function, communication module, etc. Embodiments of the present application do not limit the device form of the terminal. The terminal is usually provided with a communication module, circuit or chip for performing corresponding communication functions. The terminal can also be configured with program instructions for performing corresponding communication functions.
[0096] It should be understood that in some scenarios, the UE can also be used to act as a base station. For example, the UE can act as a scheduling entity that provides sidelink signals between UEs in V2X, D2D or point-to-point scenarios.
[0097] In embodiments of the present application, the device for implementing the functions of the terminal device, i.e. the terminal device, can be a terminal device or a device capable of supporting the terminal device to implement the functions, such as a chip system or a chip or a circuit or a communication module (i.e. a communication module for performing communication functions), which can be installed in the terminal device. In embodiments of the present application, the chip system can be composed of a chip or can include a chip and other discrete devices. In addition, the device can also be configured with program instructions for performing corresponding communication functions.
[0098] The RAN 100 and the terminal 120 can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; can also be deployed on the water surface; can also be deployed on aircraft, balloons and satellites in the air. Embodiments of the present application do not limit the scenarios in which the RAN 100 and the terminal 120 are located.
[0099] The CN 200 can be a future core network, or a 5G core network, or an evolved 5G core network. Taking the 5G core network as an example, the CN 200 includes an access and mobility management function (AMF) network element responsible for services such as mobility management and access management, a session management function (SMF) network element responsible for session management, a user plane function (UPF) network element responsible for user plane data packet routing and forwarding and quality of service (QoS) control, a policy control function (PCF) network element, and the like. The above core network network elements can work independently, or can be combined together to implement certain control functions, for example, the AMF, the SMF, and the PCF can be combined together as a core network device.
[0100] It should be understood that the above naming is only defined for the purpose of distinguishing different functions, and should not constitute any limitation on the present application. The present application does not exclude the possibility of using other names in the 5G network and other future networks. For example, in future communication networks, part or all of the above network elements can continue to use the terms in 5G, or other names can be used, etc.
[0101] It can be understood that FIG. 1 is only an example and does not constitute any limitation on the protection scope of the present application. The communication method provided by the embodiments of the present application can also involve network elements not shown in FIG. 1, and of course the communication method provided by the embodiments of the present application can also only include part of the network elements shown in FIG. 1.
[0102] For the purpose of understanding the embodiments of the present application, the terms involved in the present application are briefly explained.
[0103] 1) Artificial intelligence model
[0104] In order to support artificial intelligence (AI) technology in a wireless network, an AI node can also be introduced in the network.
[0105] The AI node can be deployed in one or more of the following positions in the communication system 10: an access network node (RAN node), a terminal device, or a core network device, etc., or the AI node can be deployed separately, for example, in a position other than any of the above devices, such as a host or a cloud server of an over the top (OTT) system. The AI node can communicate with other devices in the communication system, which can be one or more of the following: a network device, a terminal device, or a network element of a core network, etc.
[0106] It can be understood that the number of AI nodes is not limited in the present application. For example, when there are multiple AI nodes, the multiple AI nodes can be divided based on functions, such as different AI nodes being responsible for different functions.
[0107] It can also be understood that the AI nodes can be independent devices, can be integrated into the same device to implement different functions, or can be network elements in a hardware device, or can be software functions running on a dedicated hardware, or virtualized functions instantiated on a platform (for example, a cloud platform), and the specific form of the AI nodes is not limited in the present application.
[0108] The AI nodes can be AI network elements or AI modules.
[0109] Referring to FIG. 2, FIG. 2 is a schematic block diagram of AI model deployment provided by an embodiment of the present application.
[0110] As shown in FIG. 2, the network elements in the communication system are connected through interfaces (for example, NG, Xn) or air interfaces. One or more AI modules (only one is shown in FIG. 2 for clarity) are arranged in one or more of the network element nodes, such as a core network device, an access network node (RAN node), a terminal, or a device for operations administration and maintenance (OAM). The access network node can be a separate RAN node, or can include multiple RAN nodes, for example, including a CU and a DU. The CU and / or the DU can also be arranged with one or more AI modules. The CU can also be split into a CU-CP and a CU-UP, and the CU-CP and / or the CU-UP can be arranged with one or more AI modules.
[0111] The AI modules are used to implement corresponding AI functions. The AI modules deployed in different network elements can be the same or different. The model of the AI module can implement different functions according to different parameter configurations. The model of the AI module can be configured based on one or more of the following parameters: a structural parameter (for example, at least one of the number of neural network layers, the width of the neural network, the connection relationship between layers, the weight of neurons, the activation function of neurons, or the bias in the activation function), an input parameter (for example, the type of input parameters and / or the dimension of input parameters), or an output parameter (for example, the type of output parameters and / or the dimension of output parameters). The bias in the activation function can also be referred to as the bias of the neural network.
[0112] In one example, the neural network described above can be a DNN, a convolutional neuron network (CNN), a recurrent neural network (RNN), or a generative adversarial network (GAN).
[0113] A DNN is a type of artificial neural network architecture that has multiple layers of nonlinear transformation units stacked together in a hierarchical structure, forming a deep computational model. Compared to a shallow neural network, a deep neural network has more hidden layers, allowing the network model to capture more complex data intrinsic structures and high-level abstract features.
[0114] A CNN is a type of deep neural network with a convolutional structure. A CNN includes a feature extractor composed of convolutional layers and subsampling layers. The feature extractor can be viewed as a filter, and the convolution process can be viewed as using a trainable filter to convolve with an input image or a convolutional feature map.
[0115] An RNN is a type of recursive neural network that takes sequence data as input, performs recursion in the evolution direction of the sequence, and connects all nodes (recurrent units) in a chain.
[0116] A GAN is a type of deep learning model. It is composed of a generator and a discriminator, and is trained through adversarial learning. The goal is to estimate the latent distribution of data samples and generate new data samples.
[0117] An AI module can have one or more models. A model can infer an output, which includes a parameter or multiple parameters. The learning process, training process, or inference process of different models can be deployed in different nodes or devices, or can be deployed in the same node or device.
[0118] 2) Radio resource control (RRC) state
[0119] The RRC state includes RRC connected (RRC_connected), RRC idle (RRC_idle), and RRC inactive (RRC_inactive).
[0120] When the RRC is in the connected state, the terminal establishes a communication connection with the network device, and can perform data transmission and communication. In the connected state, the terminal and the network device maintain a persistent communication connection for real-time data transmission and communication services, such as voice calls, video calls, and real-time data transmission.
[0121] When the RRC is in the idle state, the terminal does not establish a connection with the network device, and the terminal is in an idle state, but still maintains a certain contact with the network device. In the idle state, the terminal periodically interacts with the network device through signaling to be ready to enter the connected state and receive a wake-up signal or data transmission at any time.
[0122] When the RRC is in the inactive state, the terminal switches from the idle state to the connected state more quickly, so that the terminal is more power-efficient and can be quickly awakened.
[0123] 3) Extended Reality (XR) Technology
[0124] XR technology includes virtual reality (VR) technology, augmented reality (AR) technology, and mixed reality (MR) technology.
[0125] VR technology refers to the rendering of visual and audio scenes to simulate the visual and audio stimulation of the real world to the user's senses as much as possible. VR technology usually requires the user to wear a head-mounted display (HMD) to completely replace the user's field of view with simulated visual components, and requires the user to wear earphones to provide audio to the user. In addition, head and motion tracking of the user in VR is also required to update the simulated visual and audio content in time to achieve consistency between the user's experience of visual and audio content and the user's actions.
[0126] AR technology refers to providing additional visual or auditory information or artificially generated content in the real environment perceived by the user. The user's perception of the real environment can be direct, that is, without intermediate perception measurement, processing and rendering, or indirect, that is, through sensors and other means for transmission and further enhancement processing.
[0127] MR technology is a further development of AR technology. This technology presents virtual scene information in a real scene to build an interactive information loop between the real world, the virtual world and the user to enhance the user's experience of reality. For example, some virtual elements are inserted into a physical scene to provide the user with an immersive experience that these virtual elements are part of the real scene.
[0128] In the XR technology, an image or a video is transmitted uplink to perform object detection and recognition on image content in the picture. At present, the object detection and recognition technology based on DNN has high detection accuracy and real-time performance. The DNN generally includes multiple neural network layers, and the calculation characteristics of different neural network layers are different, which makes the calculation amount required to be completed by each neural network layer and the data amount output by each neural network layer different. The DNN includes a convolutional layer, a pooling layer, a fully connected layer and an activation layer. When the DNN performs object detection and recognition, the neural network layers close to the input side are the convolutional layer and the pooling layer, and therefore, the calculation amount of the convolutional layer and the pooling layer is relatively small; the neural network layers close to the output side are the fully connected layer, and therefore, the calculation amount of the fully connected layer is relatively large.
[0129] At present, with the development of terminals, the terminals themselves also have certain computing capabilities and can undertake part of the neural network calculation. For example, the front-end preprocessing calculation part of the neural network is allocated to the terminal.
[0130] Referring to FIG. 3, as an example, FIG. 3 is a schematic diagram of DNN model calculation division provided by an embodiment of the present application.
[0131] As shown in FIG. 3, the DNN model includes an original image input, a terminal calculation part and a server calculation part. Specifically, in the terminal calculation part, the terminal performs preprocessing on the input original image, for example, feature extraction of the original image. Then the terminal sends the extracted features to the network device, the network device sends the features to the server, and the server calculation part performs subsequent processing on the features of the image, which can reduce the data amount processed by the server.
[0132] With the introduction of large language models, the number of layers and parameters of the DNN model are increasing, in order to reduce the calculation amount of the DNN, a network device is introduced to perform part of the calculation in the DNN to reduce the calculation burden of the DNN.
[0133] Referring to FIG. 4, as an example, FIG. 4 is a schematic diagram of network device calculation provided by an embodiment of the present application.
[0134] As shown in FIG. 4, the network device 410 includes a storage device 411 and a computing device 412, wherein the storage device 411 stores a model #1, a model #2 and a model #3, and the computing device 412 includes a graphic memory. The network device 410 serves a cell including a terminal 420, a terminal 430 and a terminal 440, and the network device 410 communicates with the terminal 420, the terminal 430 and the terminal 440, and the terminal 420, the terminal 430 and the terminal 440 use different models, for example, the terminal 420 requests the network device 410 to use the model #1 for calculation, the terminal 430 requests the network device 410 to use the model #2 for calculation, and the terminal 440 requests the network device 410 to use the model #3 for calculation. The graphic memory in the computing device 412 can also be replaced by a memory.
[0135] However, the network device needs to load the corresponding model data from the storage device to the memory in the computing device when performing a calculation task, and the deployment of the network device limits the computing capacity and the number of the network device, thus the network device needs to frequently load the model data from the storage device, which increases the running time of the calculation task; in addition, when a large amount of memory resources are configured to reduce the running time of the calculation task, resource waste is caused.
[0136] In view of this, the present application provides a communication method and a communication device, which can reduce the time of loading data from the storage device, improve the computing efficiency of the network device, and reduce resource waste.
[0137] The communication method and the communication device provided by the present application will be further described below with reference to the accompanying drawings. It can be understood that the network device and the terminal are taken as an example to illustrate the execution subject of the interaction in the present application, but the present application does not limit the execution subject of the interaction. For example, the method executed by the network device in the present application can also be implemented by a module (such as a circuit, a chip or a chip system, etc.) in the network device, or a logical node, a logical module or software capable of realizing all or part of the network function; the method executed by the terminal in the present application can also be implemented by a communication module in the terminal or a circuit or a chip (such as a modem chip (also known as a baseband chip), or a SoC chip containing a modem core, or a SIP chip) responsible for the communication function in the terminal.
[0138] Referring to FIG. 5, as an example, FIG. 5 is a schematic diagram of a communication method 500 provided by an embodiment of the present application. The method 500 shown in FIG. 5 can include the following steps.
[0139] S510, the terminal sends first request information to the network device. Correspondingly, the network device receives the first request information from the terminal.
[0140] The first request information is used to request that AI model data occupies memory resources.
[0141] For example, the AI model data can include input data, intermediate data and trained data of the AI model. The input data can be data that needs to be trained or inferred by the AI model. The intermediate data can be intermediate variables generated in the training or inference process of the AI model. The trained data is model parameters generated after the AI model is trained on the training data. The data of the AI model can also be other data that needs to be processed by the AI model, which is not limited in the present application.
[0142] For example, the memory resources can be memory resources in a central processing unit (CPU), memory resources in a GPU, or memory resources in a special chip or a special chip that may appear in the future. The special chip can be a chip specially used for data processing of the AI model, or a chip specially used for other models capable of data processing, which is not limited in the present application. For example, the memory resources can be internal storage resources of a network device, which is also called video memory. That is, the memory resources can be video memory resources in the network device, which are used to store data of a model loaded from a storage unit of the network device. For example, if a disk stores an LLaMA-2 model, the data of the LLaMA-2 model is loaded into the video memory.
[0143] It should be noted that the terminal can request that the AI model data occupies all memory resources, or request that the AI model data occupies part of the memory resources, which is not limited in the embodiments of the present application.
[0144] It should be further noted that the first request information sent by the terminal can also be received by other devices with computing capabilities, which is not limited in the embodiments of the present application. In addition, for convenience of description, the network device is taken as an example for description hereinafter.
[0145] In a possible implementation, the first request information is also used to request the network device to perform calculation using the AI model.
[0146] When the first request information is used to request the network device to perform calculation using the AI model, the AI model is an AI model corresponding to the terminal, that is, the AI model is a model used by the terminal when the terminal hopes that the network device performs a calculation task.
[0147] It should be noted that the terminal can initiate a request for using the AI model for calculation to the network device through other request information, that is, the request information for requesting the AI model data to occupy the memory resource and the request information for requesting to use the AI model for calculation can be the same request information, for example, the first request information; or can be different request information, for example, the request information for requesting to use the AI model for calculation is the third request information.
[0148] In a possible implementation, the first request information includes at least one of the following: a time length of requesting the AI model data to occupy the memory resource; a time period of requesting the AI model data to occupy the memory resource; and a size of requesting the AI model data to occupy the memory resource.
[0149] In this application, the terminal sends the first request information to the network device according to its own needs, for example, the priority of the data related to the calculation task currently processed by the terminal is high, or the number of the data related to the calculation task currently processed by the terminal is large.
[0150] As an example, taking the network device using the AI model for calculation as an example, according to the priority of the data related to the calculation task being processed, the terminal can request the AI model data to occupy the memory resource in a continuous occupation state, or can request the AI model data to occupy the memory resource in a temporary occupation state, for example, when the priority of the data related to the calculation task being processed by the terminal is high, the AI model data is requested to occupy the memory resource in a continuous occupation state; for another example, when the priority of the data related to the calculation task being processed by the terminal is low, the AI model data is requested to occupy the memory resource in a temporary occupation state.
[0151] Optionally, the first request information is specifically used to request any of the following: the AI model data to occupy the memory resource in a continuous occupation state; the AI model data to occupy the memory resource in a temporary occupation state; the AI model data to occupy the memory resource in a continuous occupation state and a time of occupying the memory resource in the continuous occupation state; the AI model data to occupy the memory resource in a temporary occupation state and a time of occupying the memory resource in the temporary occupation state.
[0152] In the continuous occupation state, the memory resource occupied by the AI model data is prohibited to be occupied by data other than the AI model data, and in the temporary occupation state, the memory resource occupied by the AI model data is allowed to be occupied by data other than the AI model data.
[0153] It should be noted that the data other than the AI model data can be understood as the terminal initiating a request to occupy the memory resource for data other than the AI model data. The data other than the AI model data can be data of other AI models or data that needs to be processed by the network device, for example, data that needs to be calculated by the network device.
[0154] It should be further noted that when the memory resource occupied by the AI model data is prohibited from being occupied by data other than the AI model data, the state of the memory resource being occupied at this time is referred to as a continuous occupation state. When the memory resource occupied by the AI model data is allowed to be occupied by data other than the AI model data, the state of the memory resource being occupied at this time is referred to as a temporary occupation state. In the application, the continuous occupation state can also be referred to as a continuous occupation state or a maintained occupation state, that is, the continuous occupation state in the embodiments of the present application can be replaced by the continuous occupation state or the maintained occupation state; the temporary occupation state can also be referred to as a non-continuous occupation state or a temporary occupation state, that is, the temporary occupation state in the embodiments of the present application can be replaced by the non-continuous occupation state or the temporary occupation state.
[0155] It can be understood that the names of the continuous occupation state and the temporary occupation state in the embodiments of the present application are not limited, as long as the state of the memory resource being occupied is consistent with the continuous occupation state described above, or the state of the memory resource being occupied is consistent with the temporary occupation state described above, both of which can be considered as the continuous occupation state or the temporary occupation state in the embodiments of the present application.
[0156] As an example, the first request information is specifically used to request the AI model data to occupy the memory resource in a continuous occupation state. At this time, the memory resource is continuously occupied by the AI model data, and when other data also needs to occupy the memory resource, the memory resource being occupied by the AI model data will not be released due to the arrival of other data, that is, the memory resource occupied by the AI model data is prohibited from being occupied by other data. For example, the first request information requests the data of AI model #1 to occupy the memory resource in a continuous occupation state. At this time, if the terminal needs to request the data of AI model #2 to occupy the memory resource, the part of the memory resource being occupied by the data of AI model #1 will not be occupied by the data of AI model #2. That is, the memory resource occupied by the data of AI model #1 will not be released to be used by the data of AI model #2 due to the arrival of the data of AI model #2, and the memory resource occupied by the data of AI model #1 is prohibited from being occupied by the data of AI model #2.
[0157] As another example, the first request information is specifically for requesting the AI model data to occupy the memory resource in a temporary occupation state, at this time, the memory resource is temporarily occupied by the AI model data, and when other data also needs to occupy the memory resource, the memory resource occupied by the AI model data is released due to the arrival of the other data, that is, the memory resource occupied by the AI model data is allowed to be occupied by the other data. For example, the first request information requests the data of AI model #1 to occupy the memory resource in a temporary occupation state, at this time, if the terminal needs to request the data of AI model #2 to occupy the memory resource, and the priority of the computing task executed by AI model #2 is higher than the priority of the computing task executed by AI model #1, the memory resource occupied by the data of AI model #1 will be occupied by the data of AI model #2, that is, the memory resource occupied by the data of AI model #1 will be released due to the arrival of the data of AI model #2 to allow the data of AI model #2 to use.
[0158] Optionally, in the continuous occupation state, the duration of the AI model data occupying the memory resource is within the first duration, and the memory resource is prohibited from being occupied by data other than the AI model data. Or in the continuous occupation state, the duration of the AI model data occupying the memory resource is less than the first duration, and the memory resource is prohibited from being occupied by data other than the AI model data.
[0159] In the continuous occupation state, the duration of the AI model data occupying the memory resource is within the first duration, and the memory resource is prohibited from being occupied by data other than the AI model data, which can be understood as follows: the duration of the AI model data occupying the memory resource is the first duration, within the first duration, other data cannot occupy the memory resource occupied by the AI model data, and after the first duration, other data can occupy the memory resource occupied by the AI model data. Or, when the duration of the AI model data occupying the memory resource is less than the first duration, other data cannot occupy the memory resource occupied by the AI model data. For example, the value of the first duration is 60 minutes, then the duration of the AI model data occupying the memory resource is 60 minutes, within the 60 minutes, the memory resource will not be occupied by data other than the AI model data.
[0160] Optionally, in the temporary occupation state, after the duration of the AI model data occupying the memory resource is greater than or equal to the second duration, the memory resource is allowed to be occupied by data other than the AI model data. Or in the temporary occupation state, after the duration of the AI model data occupying the memory resource is greater than the second duration, the memory resource is allowed to be occupied by data other than the AI model data.
[0161] The memory resource occupied by the AI model data for a time period greater than or equal to the second time period is allowed to be occupied by data other than the AI model data. It can be understood that, after the memory resource occupied by the AI model data for a time period greater than the second time period, or after the memory resource occupied by the AI model data for a time period equal to the second time period, the memory resource is allowed to be occupied by data other than the AI model data. For example, the value of the second time period is 3 minutes. After 3 minutes, if other data needs to occupy the memory resource, the memory resource currently occupied by the AI model data will be released, and the other data will occupy the memory resource. In this application, the first time period is greater than the second time period.
[0162] It should be noted that, in an implementation manner, in the temporary occupation state, the memory resource occupied by the AI model data is still occupied by other data, even if the time period of the memory resource occupied by the AI model data is not greater than the second time period, because the priority of the calculation task of calculating other data is higher than the priority of the calculation task currently executed by the AI model.
[0163] It should be further noted that the continuous occupation state and the temporary occupation state can be pre-defined by a protocol.
[0164] Further, the network device sends first indication information to the terminal according to the content requested by the first request information and in combination with its own conditions. The specific process is as follows.
[0165] S520, the network device sends first indication information to the terminal. Correspondingly, the terminal receives the first indication information from the network device.
[0166] The first indication information is used to indicate the memory resource occupation state of the AI model data. The memory resource occupation state is one of candidate occupation states, and the candidate occupation states include the continuous occupation state and the temporary occupation state.
[0167] For example, the network device sends the first indication information to the terminal in combination with the size of the memory resource of the CPU or GPU of the network device. Alternatively, when the first request information requests the network device to use the AI model for calculation, the first indication information is sent to the terminal according to the priority of the calculation task requested by the first request information. For example, the priority of the calculation task requested by the first request information is high, and the first indication information is used to indicate that the AI model data occupies the memory resource in the continuous occupation state. For another example, the storage space of the memory resource of the CPU or GPU in the network device is small, and the first indication information is used to indicate that the AI model data occupies the memory resource in the temporary occupation state.
[0168] Optionally, the first indication information is used to indicate the memory resource occupation state of the AI model data, including: the first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is a continuous occupation state; or the first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the continuous occupation state, and is specifically used to indicate a time at which the AI model data occupies the memory resource in the continuous occupation state; or the first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is a temporary occupation state; or the first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the temporary occupation state, and is specifically used to indicate a time at which the AI model data occupies the memory resource in the temporary occupation state.
[0169] The time at which the AI model data occupies the memory resource in the continuous occupation state can be understood as a time period or length of time at which the AI model data occupies the memory resource in the continuous occupation state. The time at which the AI model data occupies the memory resource in the temporary occupation state can be understood as a time period or length of time at which the AI model data occupies the memory resource in the temporary occupation state.
[0170] Optionally, the candidate occupation state further includes a prohibited occupation state, in which the AI model data is prohibited from occupying the memory resource.
[0171] Exemplarily, the first indication information is used to indicate the memory resource occupation state of the AI model data, including: the first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the prohibited occupation state.
[0172] In the first example, the first request information is specifically used to request the AI model data to occupy the memory resource in the continuous occupation state. At this time, if the priority of the computing task executed by the AI model is high or the space of the memory resource of the network device is sufficient, the first indication information sent by the network device to the terminal can indicate that the AI model data can occupy the memory resource in the continuous occupation state, i.e., the first indication information indicates that the memory resource occupation state of the AI model data is the continuous occupation state. If the priority of the computing task executed by the AI model is low or the space of the memory resource of the network device is insufficient, the first indication information sent by the network device to the terminal can indicate that the AI model data cannot occupy the memory resource in the continuous occupation state requested by the terminal, i.e., the first indication information indicates that the memory resource occupation state of the AI model data is the prohibited occupation state; or the first indication information sent by the network device to the terminal can indicate that the AI model data cannot occupy the memory resource in the continuous occupation state requested by the terminal, but indicates that the AI model data can occupy the memory resource in the temporary occupation state, i.e., the first indication information indicates that the memory resource occupation state of the AI model data is the temporary occupation state.
[0173] When the first indication information indicates that the memory resource occupation state of the AI model data is the continuous occupation state, the first indication information further indicates a time at which the AI model data occupies the memory resource in the continuous occupation state.
[0174] In the second example, the first request information specifically requests the AI model data to occupy the memory resource in the temporary occupation state. At this time, if the priority of the computing task performed by the AI model is high, the first indication information sent by the network device to the terminal can indicate that the AI model data can occupy the memory resource in the temporary occupation state, i.e., the first indication information indicates that the memory resource occupation state of the AI model data is the temporary occupation state. If the priority of the computing task performed by the AI model is low, the first indication information sent by the network device to the terminal can indicate that the AI model data cannot occupy the memory resource in the temporary occupation state requested by the terminal, i.e., the first indication information indicates that the memory resource occupation state of the AI model data is the prohibited occupation state. Or the first indication information sent by the network device to the terminal can indicate that the AI model data cannot occupy the memory resource in the temporary occupation state requested by the terminal.
[0175] When the first indication information indicates that the memory resource occupation state of the AI model data is the temporary occupation state, the first indication information further indicates a time at which the AI model data occupies the memory resource in the temporary occupation state.
[0176] As an example, when the first request information requests the AI model data to occupy the memory resource in the temporary occupation state and a time at which the AI model data occupies the memory resource in the temporary occupation state, the first indication information further indicates the time at which the AI model data occupies the memory resource in the temporary occupation state. The time at which the AI model data occupies the memory resource in the temporary occupation state requested by the first request information can be the same as or different from the time at which the AI model data occupies the memory resource in the temporary occupation state indicated by the first indication information. For example, when the memory resource of the network device is sufficient, the time at which the AI model data occupies the memory resource in the temporary occupation state requested by the first request information can be the same as the time at which the AI model data occupies the memory resource in the temporary occupation state indicated by the first indication information. For another example, when the memory resource of the network device is insufficient, the time at which the AI model data occupies the memory resource in the temporary occupation state requested by the first request information is greater than the time at which the AI model data occupies the memory resource in the temporary occupation state indicated by the first indication information.
[0177] Optionally, when the first indication information indicates that the memory resource occupation state of the AI model data is the prohibited occupation state, the first indication information further indicates a first time interval. At this time, the method 500 further includes:
[0178] The terminal sends second request information after the first time interval, the second request information being used for requesting the AI model data to occupy the memory resource.
[0179] Specifically, the network device makes a judgment of prohibiting the AI model data from occupying the memory resource in combination with its own conditions, and then sends first indication information to the terminal to indicate that the memory resource occupation state of the AI model data is a prohibited occupation state, and also indicates a first time interval to indicate the terminal to initiate a request of the AI model data occupying the memory resource again after the first time interval, that is, the terminal sends second request information.
[0180] S530, the terminal determines the memory resource occupation state of the AI model data based on the first indication information.
[0181] The detailed description of the memory resource occupation state can be referred to the above step S520, and will not be described here.
[0182] Optionally, determining the memory resource occupation state of the AI model data based on the first indication information comprises: determining, based on the first indication information, that the memory resource occupation state of the AI model data is a continuous occupation state; or determining, based on the first indication information, that the memory resource occupation state of the AI model data is a continuous occupation state and determining a time when the AI model data occupies the memory resource in the continuous occupation state; or determining, based on the first indication information, that the memory resource occupation state of the AI model data is a temporary occupation state; or determining, based on the first indication information, that the memory resource occupation state of the AI model data is a temporary occupation state and determining a time when the AI model data occupies the memory resource in the temporary occupation state.
[0183] As an example, when the first request information is specifically used for requesting the AI model data to occupy the memory resource in a continuous occupation state, if the first indication information indicates that the memory resource occupation state of the AI model data is a continuous occupation state, the terminal determines, based on the first indication information, that the memory resource occupation state of the AI model data is a continuous occupation state, at this time, the AI model data can be continuously stored in the memory resource, and the memory resource is prohibited to be occupied by other data when other data arrives. If the first indication information indicates that the memory resource occupation state of the AI model data is a continuous occupation state and a time when the AI model data occupies the memory resource in the continuous occupation state, the terminal determines, based on the first indication information, that the memory resource occupation state of the AI model data is a continuous occupation state and determines the time when the AI model data occupies the memory resource in the continuous occupation state.
[0184] As another example, when the first request information is specifically used to request the AI model data to occupy the memory resource in the temporary occupation state, if the first indication information indicates that the memory resource occupation state of the AI model data is the temporary occupation state, the terminal determines, based on the first indication information, that the memory resource occupation state of the AI model data is the temporary occupation state, at this time, the AI model data can be temporarily stored in the memory resource, and when other data with higher priority arrives, the memory resource is allowed to be occupied by the other data. If the first indication information indicates that the memory resource occupation state of the AI model data is the temporary occupation state, and the time when the memory resource is occupied by the AI model data in the temporary occupation state, the terminal determines, based on the first indication information, that the memory resource occupation state of the AI model data is the temporary occupation state, and determines the time when the memory resource is occupied by the AI model data in the temporary occupation state.
[0185] As another example, if the first indication information indicates that the memory resource occupation state of the AI model data is the prohibited occupation state, the terminal determines, based on the first indication information, that the memory resource occupation state of the AI model data is the prohibited occupation state. Specifically, if the first request information is specifically used to request the AI model data to occupy the memory resource in the continuous occupation state, the network device prohibits the AI model data from occupying the memory resource in the continuous occupation state requested by the terminal. If the first request information is specifically used to request the AI model data to occupy the memory resource in the temporary occupation state, the network device prohibits the AI model data from occupying the memory resource in the temporary occupation state requested by the terminal.
[0186] Optionally, when the terminal is in a radio resource control (RRC) connected state, the first request information is specifically used to request the AI model data to occupy the memory resource in the continuous occupation state. That is, there is an association relationship between the RRC connected state and the continuous occupation state, and when the terminal is in the RRC connected state, the AI model data can only occupy the memory resource in the continuous occupation state.
[0187] Optionally, when the terminal is in an RRC connected state or an RRC inactive state, the first request information is specifically used to request the AI model data to occupy the memory resource in the temporary occupation state. That is, there is an association relationship between the RRC connected state or the RRC inactive state and the temporary occupation state, and when the terminal is in the RRC connected state, the AI model data can only occupy the memory resource in the temporary occupation state.
[0188] Based on the above association relationship between the RRC state and the continuous occupation state or the temporary occupation state, when the RRC state is switched, the memory resource occupation state will also change. For example, if the memory resource occupation state is the continuous occupation state at this time, when the terminal device is switched from the RRC connected state to the RRC inactive state or the RRC idle state, the AI model data will no longer occupy the memory resource in the continuous occupation state.
[0189] Referring to FIG. 6, as an example, FIG. 6 is a schematic diagram of RRC state switching provided by an embodiment of the present application.
[0190] As shown in FIG. 6, the RRC state in which the terminal is located is switched between the RRC connected state, the RRC idle state and the RRC inactive state. For example, under case 1 or 2, the terminal is switched from the RRC connected state to the RRC inactive state, at this time, the AI model data will no longer occupy the memory resource in the persistent occupation state, or the AI model data will no longer occupy the memory resource in the temporary occupation state; under case 3, the terminal is switched from the RRC inactive state to the RRC idle state, at this time, the AI model data will no longer occupy the memory resource in the temporary occupation state; under case 4 or 5, the terminal device is switched from the RRC idle state to the RRC connected state, at this time, the terminal needs to re-initiate the request information (for example, the first request information) to the network device to request the AI model data to occupy the memory resource, for example, to request the AI model data to occupy the memory resource in the persistent occupation state, or to request to occupy the memory resource in the temporary occupation state; under case 6, the terminal device is switched from the RRC inactive state to the RRC connected state, at this time, the terminal needs to re-initiate the request information (for example, the first request information) to the network device to request the AI model data to occupy the memory resource in the persistent occupation state.
[0191] In this way, by associating the RRC state of the terminal with the occupation state of the memory resource, when the RRC connected state of the terminal is converted, the memory resource will not be in the state of being occupied by the AI model data requested by the terminal when the terminal device is in the RRC inactive state or the idle state, and when the RRC state changes, the occupation state of the memory resource also changes, thereby improving the utilization rate of the memory resource.
[0192] In the embodiment of the present application, the terminal initiates a request for the AI model data to occupy the memory resource, and receives the indication information, based on which the terminal can determine the memory resource occupation state of the AI model data, which can be one of the persistent occupation state and the temporary occupation state, so that for the memory resource occupied by the AI model data in the persistent occupation state, other data is prohibited from occupying the memory resource, reducing the time of loading data from the storage device and reducing the running time of the computing task, thereby improving the computing efficiency of the network device; and for the memory resource occupied by the AI model data in the temporary occupation state, other data is allowed to occupy the memory resource, thereby reducing the waste of resources.
[0193] The method provided in the embodiments of the present application is described in detail above in combination with FIG. 5 to FIG. 6. The apparatus provided in the embodiments of the present application is described in detail below in combination with FIG. 7 to FIG. 8. It should be understood that the description of the apparatus embodiments corresponds to the description of the method embodiments, and therefore, the content not described in detail can be referred to the method embodiments described above, which will not be described here for brevity.
[0194] Referring to FIG. 7, as an example, FIG. 7 is a possible exemplary block diagram of the communication apparatus involved in the embodiments of the present application. As shown in FIG. 7, the communication apparatus 1000 can include modules or units for implementing the corresponding method embodiments described above. In a possible design, the communication apparatus 1000 includes a communication unit 1003 and a processing unit 1002. Optionally, the communication apparatus 1000 can further include a storage unit 1001 for storing apparatus program codes and / or data. The communication unit 1003 can also be referred to as a communication interface, a transceiver unit or an interface unit.
[0195] The communication apparatus 1000 can be a terminal-side apparatus in the embodiments described above, for example, a terminal or a communication module in the terminal, or a circuit or chip responsible for the communication function in the terminal.
[0196] For example, in an embodiment, the processing unit 1002 is configured to determine, based on the first indication information, a memory resource occupation state of the AI model data, the memory resource occupation state being one of candidate occupation states, the candidate occupation states including a persistent occupation state and a temporary occupation state, in the persistent occupation state, the memory resource occupied by the AI model data is prohibited to be occupied by data other than the AI model data, and in the temporary occupation state, the memory resource occupied by the AI model data is allowed to be occupied by data other than the AI model data. The communication unit 1003 is configured to send first request information for requesting an artificial intelligence (AI) model data to occupy a memory resource, and the communication unit 1003 is further configured to receive first indication information.
[0197] In a possible design, in the persistent occupation state, a time length for which the memory resource is occupied by the AI model data is within a first time length, and the memory resource is prohibited to be occupied by data other than the AI model data.
[0198] In a possible design, in the persistent occupation state, the time length for which the memory resource is occupied by the AI model data is less than the first time length, and the memory resource is prohibited to be occupied by data other than the AI model data.
[0199] In a possible design, in the temporary occupation state, the time length for which the memory resource is occupied by the AI model data is greater than or equal to a second time length, and the memory resource is allowed to be occupied by data other than the AI model data.
[0200] In a possible design, in the temporary occupation state, the memory resource is allowed to be occupied by data other than the AI model data after the AI model data occupies the memory resource for a time period longer than a second time period.
[0201] In a possible design, the first request information is specifically used to request any of the following: the AI model data to occupy the memory resource in the persistent occupation state; the AI model data to occupy the memory resource in the temporary occupation state; the AI model data to occupy the memory resource in the persistent occupation state and a time period for occupying the memory resource in the persistent occupation state; the AI model data to occupy the memory resource in the temporary occupation state and a time period for occupying the memory resource in the temporary occupation state.
[0202] In a possible design, the processing unit 1002 is further configured to determine, based on the first indication information, that the memory resource occupation state of the AI model data is the persistent occupation state, or determine, based on the first indication information, that the memory resource occupation state of the AI model data is the persistent occupation state and a time period for occupying the memory resource in the persistent occupation state, or determine, based on the first indication information, that the memory resource occupation state of the AI model data is the temporary occupation state, or determine, based on the first indication information, that the memory resource occupation state of the AI model data is the temporary occupation state and a time period for occupying the memory resource in the temporary occupation state.
[0203] In a possible design, the candidate occupation state further includes a prohibited occupation state in which the AI model data is prohibited from occupying the memory resource.
[0204] In a possible design, the processing unit 1002 is further configured to determine, based on the first indication information, that the memory resource occupation state of the AI model data is the prohibited occupation state.
[0205] In a possible design, the processing unit 1002 is further configured to determine, based on the first indication information, a first time interval, and the communication unit 1003 is further configured to send, after the first time interval, second request information used to request the AI model data to occupy the memory resource.
[0206] In a possible design, when the terminal is in a radio resource control (RRC) connected state, the first request information is specifically used to request the AI model data to occupy the memory resource in the persistent occupation state.
[0207] In a possible design, when the terminal is in an RRC connected state or an RRC inactive state, the first request information is specifically used to request the AI model data to occupy the memory resource in the temporary occupation state.
[0208] In a possible design, when the communication apparatus 1000 is a terminal or a communication module in a terminal, the function of the processing unit 1002 can be implemented by one or more processors. Specifically, the processor can include a modem chip, or a system on chip (SoC) chip or a SIP chip including a modem core. The function of the communication unit 1003 can be implemented by a transceiver circuit.
[0209] In a possible design, when the communication apparatus 1000 is a circuit or chip responsible for communication functions in a terminal, such as a modem chip or a system on chip (SoC) chip or a SIP chip including a modem core, the function of the processing unit 1002 can be implemented by circuitry including one or more processors or processor cores in the chip. The function of the communication unit 1003 can be implemented by an interface circuit or a data transceiver circuit on the chip.
[0210] The communication apparatus 1000 can be a network side device in the above-described embodiments, for example, a network device, or a module (for example, a circuit, a chip or a chip system, etc.) in a network device, or a logic node or a logic module capable of implementing all or part of the function of a network device.
[0211] For example, in an embodiment, the communication unit 1003 is configured to receive first request information, where the first request information is used to request memory resource occupation of artificial intelligence (AI) model data; and the communication unit 1003 is further configured to send first indication information, where the first indication information is used to indicate a memory resource occupation state of the AI model data, and the memory resource occupation state is one of candidate occupation states, and the candidate occupation states include a persistent occupation state and a temporary occupation state. In the persistent occupation state, the memory resource occupied by the AI model data is prohibited to be occupied by data other than the AI model data. In the temporary occupation state, the memory resource occupied by the AI model data is allowed to be occupied by data other than the AI model data.
[0212] In a possible design, in the persistent occupation state, the memory resource occupied by the AI model data is prohibited to be occupied by data other than the AI model data within a first time length.
[0213] In a possible design, in the persistent occupation state, the memory resource occupied by the AI model data is prohibited to be occupied by data other than the AI model data within a time length less than the first time length.
[0214] In a possible design, in the temporary occupation state, the memory resource occupied by the AI model data is allowed to be occupied by data other than the AI model data after a time length occupied by the AI model data is greater than or equal to a second time length.
[0215] In a possible design, after the AI model data occupies the memory resource for a time period longer than a second time period in the temporary occupation state, the memory resource is allowed to be occupied by data other than the AI model data.
[0216] In a possible design, the first request information is specifically used to request any of the following: the AI model data to occupy the memory resource in the persistent occupation state; the AI model data to occupy the memory resource in the temporary occupation state; the AI model data to occupy the memory resource in the persistent occupation state and a time period for occupying the memory resource in the persistent occupation state; and the AI model data to occupy the memory resource in the temporary occupation state and a time period for occupying the memory resource in the temporary occupation state.
[0217] In a possible design, the first indication information is used to indicate the memory resource occupation state of the AI model data, and includes the following.
[0218] The first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the persistent occupation state.
[0219] The first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the persistent occupation state, and is specifically used to indicate a time period for occupying the memory resource in the persistent occupation state.
[0220] The first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the temporary occupation state.
[0221] The first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the temporary occupation state, and is specifically used to indicate a time period for occupying the memory resource in the temporary occupation state.
[0222] In a possible design, the candidate occupation states further include a prohibited occupation state, in which the AI model data is prohibited from occupying the memory resource.
[0223] In a possible design, the first indication information is used to indicate the memory resource occupation state of the AI model data, and includes the following: the first indication information is specifically used to indicate that the memory resource occupation state of the AI model data is the prohibited occupation state.
[0224] In a possible design, the first indication information is further used to indicate a first time interval, and the method further includes: receiving second request information after the first time interval, where the second request information is used to request the AI model data to occupy the memory resource.
[0225] In a possible design, the first request information is specifically used for requesting the AI model data to occupy the memory resource in the persistent occupation state when the terminal is in a radio resource control (RRC) connected state.
[0226] In a possible design, the first request information is specifically used for requesting the AI model data to occupy the memory resource in the temporary occupation state when the terminal is in an RRC connected state or an RRC inactive state.
[0227] In a possible design, when the communication apparatus 1000 is a network device or a communication module in a network device, the function of the processing unit 1002 can be implemented by one or more processors. Specifically, the processor can include a chip. The function of the communication unit 1003 can be implemented by a transceiver circuit.
[0228] In a possible design, when the communication apparatus 1000 is a circuit or a chip responsible for communication functions in a network device, the function of the processing unit 1002 can be implemented by a circuit system including one or more processors or processor cores in the chip. The function of the communication unit 1003 can be implemented by an interface circuit or a data transceiver circuit on the chip.
[0229] It can be understood that the division of the units in the apparatuses described above is merely a logical function division, one function unit can correspond to one function, or two or more functions can be integrated into one function unit. In actual implementation, all or part of the units can be integrated into one physical entity, or distributed on different physical entities. In addition, the function units described above can be implemented in the form of hardware, software, or a combination of hardware and software. Whether a function is implemented in hardware or software depends on a specific application and design constraint condition of the technical solution. A person skilled in the art can implement the described functions by using different methods for specific applications, but such implementation should not be considered beyond the scope of the present application.
[0230] In one example, the function units in any of the apparatuses described above can be one or more integrated circuits configured to implement the above methods, for example, one or more ASICs, or one or more CPUs, one or more MPUs, one or more MCUs, one or more DSPs, or one or more FPGAs, or a combination of at least two of the integrated circuit forms.
[0231] In one example, the storage unit 1001 can include random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory and / or registers, etc.
[0232] Referring to FIG. 8, as an example, FIG. 8 is a structural schematic diagram of a terminal 2000 provided by an embodiment of the present application, which can correspond to the terminal shown in FIG. 1, for implementing the operation of the terminal in the above embodiments. As shown in (a) of FIG. 8, the terminal 2000 includes one or more antennas 2010, a radio frequency processing system 2020, and a processor system 2030.
[0233] In the downlink or sidelink direction, the radio frequency processing system 2020 receives radio frequency signals through the antenna 2010 and sends the signals processed by radio frequency to the processor system 2030 for further processing. In the uplink or sidelink direction, the processor system 2030 sends the terminal-side information after signal processing to the radio frequency processing system 2020, which processes the signal by radio frequency and transmits it through the antenna 2010.
[0234] In one example, the radio frequency processing system 2020, as a communication interface for external communication of the terminal, can include a radio frequency front end 2021 (RFFE) and a radio frequency transceiver 2022. The RFFE 2021 is mainly used for one or more of shaping, passband selection, or gain processing of RF signals received by the antenna or RF signals to be sent through the antenna, and can include one or more of radio frequency switches, duplexers, filters, power amplifiers, antenna tuning, and low-noise amplifiers. The RFFE 2021 can be a circuit system composed of multiple discrete devices, or can be integrated and packaged in one or more chips. The radio frequency transceiver 2022 is used to process the RF signals received by the RFFE into baseband / intermediate frequency signals for further processing by the processor system 2030, and to process the baseband / intermediate frequency signals provided by the processor system 2030 into RF signals for sending to the RFFE 2021. The baseband / intermediate frequency signals transmitted between the radio frequency transceiver 2022 and the processor system 2030 can be digital signals or analog signals. The radio frequency transceiver 2022 can be implemented by one or more chips, which are usually referred to as radio frequency chips (RFIC).
[0235] In one example, the processor system 2030 can include one or more processors for processing signals and for executing appropriate instructions to carry out one or more communication protocols. Optionally, the processor system 2030 can further include a memory 2036. In one example, the one or more processors include at least one baseband processor 2031 (also referred to as a modem processor). The memory 2036 is used for storing data and / or computer program instructions. Optionally, the processor system 2030 can further include one or more application processors 2032 for implementing processing for an operating system of the terminal and for application layers. Optionally, the processor system 2030 can further include one or more of a voice subsystem 2033, a multimedia subsystem 2034, or an interface circuit 2035. The voice subsystem 2033 is used for processing voice signals, the multimedia subsystem 2034 is used for processing multimedia related operations, such as video codec, image processing, etc., and the interface circuit 2035 is used for communicating with other terminal components, such as a display 2040, an input device 2050, a memory 2060, etc. The above components in the processor system 2030 can communicate with each other through a bus or a communication interface circuit.
[0236] In one example, the processor system 2030 can be packaged as a processor chip, such as a SoC chip or a SIP chip. In one example, the processor system 2030 can be a system of multiple chips, for example, the baseband processor 2031 can be packaged separately as a chip, or packaged with part or all of the circuitry of the radio frequency processing system as a chip.
[0237] In one example, the memory 2036 can be an on-chip memory, i.e., located on the chip of the processor system 2030. In one example, the memory 2060 can be an off-chip memory, i.e., located off the chip of the processor system 2030.
[0238] In one example, as shown in (b) of FIG. 8, the baseband processor 2031 in the terminal 2000 provided by the embodiments of the present application can include one or more processor cores 20311 and interface circuit 20314. The one or more processor cores 20311 are configured to process signals and execute one or more communication protocols. Optionally, the baseband processor 2031 can further include a memory 20312 configured to store at least part of corresponding computer program instructions and / or data. In one example, the one or more processor cores 20311 implement the related operations in the above method embodiments by executing the computer program instructions stored in the memory 20312. In the present application, the memory 20312 configured to store corresponding computer program instructions and / or data can mean that the memory 20312 is configured to store all corresponding computer program instructions and / or data for execution by the processor core 20311; or can mean that the memory 20312 is configured to store part of corresponding computer program instructions and / or data, which includes computer program instructions and / or data currently needed for execution by the processor core 20311, and the memory 20312 can store different parts of computer program instructions and / or data for execution by the processor core 20311 multiple times to implement the related operations in the above method embodiments. The interface circuit 20314 serves as a communication interface to realize communication with other components, such as transmitting signals with the radio frequency processing system 2020, communicating with other subsystems and related components of the processor system 2030 through a bus, such as transmitting data control signals with the application processor 2032, and transmitting data or computer program instructions with the memory 2036 or the memory 2060. Optionally, in order to reduce the load of the processor core, a baseband signal processing circuit 20313 can also be provided to implement at least part of the processing work of the baseband signal, including one or more of demodulation, modulation, encoding or decoding of signals.
[0239] In one example, the communication apparatus provided by the embodiments of the present application can be a terminal 2000, a communication module including a processor system 2030 and a radio frequency system 2020, the processor system 2030, or the baseband processor 2031.
[0240] The above processor, processor system, application processor, baseband processor, processor circuit or processor core can be collectively referred to as a processor, which can include one or more combinations of CPU, DSP, MPU, MCU, GPU, FPGA, ASIC, AI processor or NPU.
[0241] The above memory can include one or more of the following storage media: random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), phase-change memory (PCM), resistive RAM (ReRAM), magnetoresistive RAM (MRAM), ferroelectric RAM (FRAM), cache, register, read-only memory (ROM), flash memory, erasable programmable ROM (EPROM), hard disk, etc. In an example, computer program instructions for implementing the above embodiments can be stored on a non-volatile memory, such as at least part of the above memory 2060 (which can be one or more of ROM, flash memory, EPROM, or hard disk). During terminal operation, the corresponding computer program instructions can be loaded in whole or in part into a memory with faster transmission speed to the processor, such as at least part of the above memory 2036 and / or memory 20312 (which can be one or more of RAM, SRAM, DRAM, PCM, RERAM, MRAM, FRAM, cache, or register), for execution by the processor to implement the steps in the above method embodiments.
[0242] In an example, the radio frequency transceiver 2022 and the radio frequency front end 2021 can also be packaged in one chip. In an example, the radio frequency transceiver 2022, the radio frequency front end 2021, and the baseband processor 2031 can also be packaged in one chip.
[0243] The embodiments of the present application also provide a computer readable storage medium having stored thereon computer programs or instructions for implementing the methods performed by the communication device (such as a terminal, and also such as a network device) in the above method embodiments. For example, the computer programs or instructions, when running on the communication device, cause the communication device (such as a terminal, and also such as a network device) to perform the above method (such as method 500).
[0244] The embodiments of the present application further provide a computer program product comprising instructions which, when executed by a computer, implement the method performed by the communication device (e.g., the terminal, or the network device) in any of the above method embodiments. For example, when the computer program or instructions are run on the communication device, the communication device (e.g., the terminal, or the network device) performs the above method (e.g., the method 500).
[0245] The embodiments of the present application further provide a communication system comprising the terminal and / or the network device in any of the above embodiments. For example, the system comprises the terminal and the network device in the embodiment of FIG. 5.
[0246] The above-provided explanation and beneficial effects of the related content in any of the above-provided apparatuses can refer to the corresponding method embodiments provided above, and will not be repeated here.
[0247] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other means. For example, the apparatus embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or components shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.
[0248] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable apparatus. For example, the computer can be a personal computer, a server, a network device, or the like. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media sets. The available media can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD), etc. For example, the foregoing available media includes but is not limited to: a variety of media that can store program codes such as a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk, etc.
[0249] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A communication method characterized by comprising: Applied to a terminal side, the method comprises: sending first request information, the first request information being used for requesting an artificial intelligence (AI) model data to occupy memory resources; receiving first indication information, determining, based on the first indication information, a memory resource occupation state of the AI model data, the memory resource occupation state being one of candidate occupation states, the candidate occupation states comprising a persistent occupation state and a temporary occupation state, in the persistent occupation state, the memory resources occupied by the AI model data are prohibited from being occupied by data other than the AI model data, in the temporary occupation state, the memory resources occupied by the AI model data are allowed to be occupied by data other than the AI model data.
2. The method of claim 1, wherein, in the persistent occupation state, the AI model data occupies the memory resources for a time length within a first time length, and the memory resources are prohibited from being occupied by data other than the AI model data.
3. The method according to claim 1 or 2, characterized in that, in the temporary occupation state, after the AI model data occupies the memory resources for a time length greater than or equal to a second time length, the memory resources are allowed to be occupied by data other than the AI model data.
4. The method according to any one of claims 1-3, characterized in that, The first request information is specifically used for requesting any one of the following: the AI model data occupies the memory resources in the persistent occupation state; the AI model data occupies the memory resources in the temporary occupation state; the AI model data occupies the memory resources in the persistent occupation state, and a time length of occupying the memory resources in the persistent occupation state; the AI model data occupies the memory resources in the temporary occupation state, and a time length of occupying the memory resources in the temporary occupation state.
5. The method according to any one of claims 1-4, characterized in that, Determining, based on the first indication information, the memory resource occupation state of the AI model data comprises: determining, based on the first indication information, the memory resource occupation state of the AI model data as the persistent occupation state; or determining, based on the first indication information, the memory resource occupation state of the AI model data as the persistent occupation state, and determining a time length of occupying the memory resources in the persistent occupation state; or determining, based on the first indication information, the memory resource occupation state of the AI model data as the temporary occupation state; or determining, based on the first indication information, the memory resource occupation state of the AI model data as the temporary occupation state, and determining a time length of occupying the memory resources in the temporary occupation state.
6. The method of claim 1, wherein, The candidate occupation states further comprise a prohibited occupation state, in which the AI model data is prohibited from occupying the memory resources.
7. The method of claim 6, wherein, Determining, based on the first indication information, the memory resource occupation state of the AI model data comprises: determining, based on the first indication information, the memory resource occupation state of the AI model data as the prohibited occupation state.
8. The method according to any one of claims 1 to 7, characterized in that, The method further comprises: determining a first time interval based on the first indication information; sending second request information after the first time interval, the second request information being used for requesting the AI model data to occupy the memory resources.
9. The method of any one of claims 1-8, wherein The first request information is specifically used for requesting the AI model data to occupy the memory resource in the continuous occupation state when the terminal is in a radio resource control (RRC) connected state.
10. The method of any one of claims 1-9, wherein, The first request information is specifically used for requesting the AI model data to occupy the memory resource in the temporary occupation state when the terminal is in an RRC connected state or an RRC inactive state.
11. A communications device, characterized by The computer readable storage medium has stored thereon computer programs or instructions, which, when executed, cause the method of any one of claims 1-10 to be performed.
12. A computer-readable storage medium, characterized in that, The computer readable storage medium has stored thereon computer programs or instructions, which, when executed, cause the method of any one of claims 1-10 to be performed.
13. A computer program product, characterised in that, The computer program product includes computer programs or instructions, which, when executed, cause the method of any one of claims 1-10 to be performed.
Citation Information
Patent Citations
QoS guarantee method and device applied to MSNET network
CN103327542A
Resource Processing Method, Apparatus, and Medium
US20230319856A1
Method and apparatus for temporarily preempting resources, and network device
WO2022126465A1
Transfer of artificial intelligence network management models in wireless communication systems
WO2024103543A1