Inference method and communication device
The second prompt word is obtained through the first AI model to supplement context information and inferred in concert with the second AI model, solving the problem of low inference accuracy when describing in incomplete problems is achieved, achieving more efficient and accurate replies, and improving user experience.
Patent Information
- Application Number
- CN202311600465.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing AI model provides incomplete problem descriptions with low inference accuracy, resulting in poor user experience, and requires multiple interactions to obtain satisfactory responses, increasing latency.
The first prompt word of the user is received through the first AI model, and the second prompt word is obtained based on the prompt word and the local information of the first network device, supplementing the description context information. The first AI model then sends a third prompt word, including the first and second prompt words, to the second AI model, and performs collaborative reasoning to improve the accuracy of the reply.
It reduces the number of interactions with users, reduces the inference delay, improves the inference accuracy of the AI model, and improves the user experience.
Smart Images

Figure CN120046722A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and more particularly, to a method of inference and a communication device. Background Art
[0002] Artificial intelligence (AI) has been widely applied in the field of communication technologies. Currently, when a user asks a question based on some AI models (such as large language models (LLMs)), the user may not provide a comprehensive description of the question. Then, when the question prompt is relatively short, the inference accuracy of the AI model is low. If precise services are to be achieved, the AI model needs to clearly locate the user's question. Therefore, the user needs to interact with the AI model multiple times, resulting in a large delay for the AI model to infer a satisfactory answer for the user, and ultimately leading to a poor user experience. Summary of the Invention
[0003] This application provides a method of inference, enabling the AI model to provide more precise answers for users, thereby improving the user experience.
[0004] In a first aspect, a method of inference is provided. This method can be executed by a first AI model, or alternatively, by a component (such as a chip or a circuit) of the first AI model, and this is not limited.
[0005] The method includes: the first AI model receives a first prompt from a terminal device, where the first prompt is used to request a corresponding answer based on the first prompt; the first AI model obtains a second prompt based on the first prompt and local information of a first network device, where the second prompt is used to describe information corresponding to at least one parameter related to the first prompt in the local information of the first network device, and the local information of the first network device is characteristic information of a cell covered by the first network device, and the first network device is a network device in a geographical range related to the first prompt; the first AI model sends a third prompt to a second AI model, where the second AI model is an AI model for inferring the answer corresponding to the first prompt, and the third prompt includes the first prompt and the second prompt, or the first AI model sends a first answer result to the terminal device based on the third prompt.
[0006] Based on the above technical solution, the first AI model can supplement and describe the context information corresponding to the first prompt based on the local information of the first network device. Then, the first AI model can use the prompt obtained after supplementing the context information for inference, or assist the second AI model in inference. This can not only reduce the number of interactions with the user, reduce the inference delay, but also improve the inference accuracy of the AI model, ensure the user's demand for service quality, and enhance the user experience.
[0007] In combination with the first aspect, in certain implementations of the first aspect, the first AI model obtains a fourth prompt word based on the first prompt word, and the fourth prompt word is used to describe at least one parameter related to the first prompt word.
[0008] In the above technical solution, the first AI model takes the received first prompt word as the input of the AI model, and can dig out the deeper needs behind the user's prompt word. For example, if the first prompt word is "Please recommend a library", the first AI model can infer from the "library" in the first prompt word that the user will mainly move indoors and may use wireless networks to perform operations such as online access and retrieval of books, in-library navigation, and electronic cloud notes. Therefore, the fourth prompt word "Also consider the wireless service quality" can be derived and generated.
[0009] In combination with the first aspect, in certain implementations of the first aspect, the third prompt word further includes the fourth prompt word.
[0010] In combination with the first aspect, in certain implementations of the first aspect, if the first prompt word does not include geographical range information or the geographical range information included in the first prompt word does not meet the geographical range accuracy required for reasoning, before the first AI model obtains the second prompt word, the method further includes: the first AI model sends a fifth prompt word to the terminal device, and the fifth prompt word is used to prompt the input of geographical range information related to the first prompt word, or, the fifth prompt word includes a sixth prompt word and multiple alternative geographical range information, and the sixth prompt word is used to prompt to select the geographical range information related to the first prompt word from the multiple alternative geographical range information; the first AI model receives a seventh prompt word from the terminal device, and the seventh prompt word is used to indicate the geographical range information related to the first prompt word, and the first AI model obtains the local information of the first network device based on the seventh prompt word.
[0011] Based on the above technical solution, the first AI model can obtain the local data of the network device corresponding to the corresponding area based on the seventh prompt word, and supplement the context information of the first prompt word based on the obtained local data of the network device.
[0012] In combination with the first aspect, in certain implementations of the first aspect, the first AI model sends the fourth prompt word to the second AI model, and the method further includes: the first AI model receives a second reply result from the second AI model, and the second reply result is the reply result inferred by the second model based on the third prompt word; the first AI model verifies that the second reply result is a reasonable reply result based on the local information of the first network device; the first AI model sends the second reply result to the terminal device.
[0013] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the first AI model receives a third response result from the second AI model, where the third response result is a response result inferred by the second AI model based on a third prompt; the first AI model verifies that the third response result is an unreasonable response result based on the local information of the first network device, and the first AI model sends an eighth prompt to the second AI model, where the eighth prompt includes the third prompt and a ninth prompt, and the ninth prompt is used to prompt to re-infer the response result based on the third prompt.
[0014] In the above technical solution, the first AI model can use the obtained local information of the first network device to verify the inference result of the second AI model, ensuring that the response finally sent to the user is a reasonable response, and avoiding the reduction of the user experience due to the second AI model obtaining an inappropriate response result due to the hallucination phenomenon.
[0015] In combination with the first aspect, in some implementations of the first aspect, before the first AI model sends a first result to the terminal device based on the third prompt, the method further includes: the first AI model verifies that the first result is a reasonable response result based on the local information of the first network device.
[0016] In combination with the first aspect, in some implementations of the first aspect, before the first AI model sends a first response result to the terminal device based on the third prompt, the method further includes: the first AI model infers a fourth response result based on the third prompt; the first AI model verifies that the fourth response result is an unreasonable response result based on the local information of the first network device.
[0017] In the above technical solution, the first AI model can use the obtained local information of the first network device to verify its own inference result, ensuring that the response finally sent to the user is a reasonable response, and avoiding the reduction of the user experience due to the first AI model obtaining an inappropriate response result due to the hallucination phenomenon.
[0018] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the first AI model saves a first sample in the local sample library, where the first sample includes the third prompt and the final response result of the third prompt.
[0019] In combination with the first aspect, in some implementations of the first aspect, the first AI model obtains a second prompt based on the first prompt and the local information of the first network device, including: the first AI model obtains the second prompt based on the first prompt, the local information of the first network device, and a tenth prompt, where the tenth prompt is used to describe at least one sample in the sample library related to the first prompt.
[0020] In the above solution, historical examples of the user are added before the first prompt word, enriching the context information of the user's prompt word based on the user's historical examples. This enables the AI model for the inference reply result to also provide precise personalized services for the user based on information related to the user's habits or preferences during inference, thereby further enhancing the user experience.
[0021] Combined with the first aspect, in some implementation manners of the first aspect, the first AI model is the AI model for inferring the reply corresponding to the first prompt word. The first AI model is the AI model of the cloud or the network device of the service terminal device. Or, the second AI model is the AI model for inferring the reply corresponding to the first prompt word. The first AI model is the AI model of the network device of the service terminal device, and the second AI model is the AI model of the cloud.
[0022] Exemplarily, the first AI model and the second AI model are large language models (LLMs).
[0023] In a second aspect, a communication device is provided. This device is used to execute the method provided in the first aspect above. Specifically, the device may include units and / or modules for executing the method in any aspect or any possible implementation manner in the first aspect, such as a processing unit and / or a communication unit.
[0024] In one implementation manner, the device is the first AI model. When the device is the first AI model, the communication unit may be a transceiver or an input / output interface; the processing unit may be at least one processor. Optionally, the transceiver may be a transceiver circuit. Optionally, the input / output interface may be an input / output circuit.
[0025] In another implementation manner, the device is a chip, a chip system, or a circuit for the first AI model. When the device is a chip, a chip system, or a circuit for the first AI model, the communication unit may be an input / output interface, an interface circuit, an output circuit, an input circuit, a pin, or a related circuit, etc. on the chip, the chip system, or the circuit; the processing unit may be at least one processor, a processing circuit, or a logic circuit, etc.
[0026] In a third aspect, a communication device is provided. The device includes: at least one processor, at least one processor is coupled to at least one memory, at least one memory is used to store a computer program or instruction, and at least one processor is used to call and run the computer program or instruction from at least one memory, so that the communication device executes the method in the first aspect or any possible implementation manner in the first aspect.
[0027] In one implementation manner, the device is the first AI model.
[0028] In another implementation, the device is a chip, a chip system, or a circuit used in the first AI model.
[0029] In a fourth aspect, a processor is provided for executing the method provided in the first aspect.
[0030] For operations such as sending and obtaining / receiving involved in the processor, if there is no special description, or if it does not conflict with its actual role or internal logic in the relevant description, it can be understood as the operations of the processor for outputting and receiving, inputting, etc., and can also be understood as the sending and receiving operations performed by the radio frequency circuit and the antenna. This application does not make any limitations in this regard.
[0031] In a fifth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores program codes for a device to execute, and the program codes include methods for executing the above-mentioned first aspect or any possible implementation manner in the first aspect.
[0032] In a sixth aspect, a computer program product containing instructions is provided. When the computer program product runs on a computer, the computer is caused to execute the method in the above-mentioned first aspect and any possible implementation manner in the first aspect.
[0033] In a seventh aspect, a chip is provided. The chip includes a processor and a communication interface. The processor reads instructions stored on a memory through the communication interface and executes the method in the above-mentioned first aspect or any possible implementation manner in the first aspect.
[0034] Optionally, as an implementation manner, the chip further includes a memory. A computer program or instructions are stored in the memory, and the processor is used to execute the computer program or instructions stored on the memory. When the computer program or instructions are executed, the processor is used to execute the method in the above-mentioned first aspect or any possible implementation manner in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a schematic diagram of an application framework 100 in a communication system applicable to an embodiment of the present application.
[0036] Figure 2 is a schematic diagram of an application framework 200 in a communication system applicable to an embodiment of the present application.
[0037] Figure 3 is a schematic diagram of a communication system 300 applicable to an embodiment of the present application.
[0038] Figure 4 is a schematic diagram of a neuron structure.
[0039] Figure 5 is a schematic flowchart of a communication method 500 provided by the present application.
[0040] Figure 6 It is a schematic diagram of a method 600 for cloud-network collaborative inference proposed in this application.
[0041] Figure 7 It is a schematic structural diagram of an AI model of a possible network device.
[0042] Figure 8 It is a schematic block diagram of a communication device 800 provided by an embodiment of this application.
[0043] Figure 9 It is a schematic block diagram of a communication device 900 provided by an embodiment of this application. Detailed implementation manners
[0044] Next, the technical solutions in the embodiments of this application will be described with reference to the accompanying drawings.
[0045] Before introducing the embodiments of this application, the following points are first explained.
[0046] 1. In the description of the embodiments of this application, unless otherwise specified, "a plurality of" means two or more.
[0047] 2. In each embodiment of this application, if there is no special specification and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be mutually referred to, and the technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.
[0048] 3. The various numerical numbers involved in this application are only for the convenience of description and do not limit the scope of this application. The size of the serial numbers involved in this application does not mean the order of execution. The execution order of each process should be determined by its function and internal logic. For example, the terms "first", "second", "third", "fourth" and other various term numbers (if any) in the specification, claims and drawings of this application are used to distinguish similar objects and do not limit the size, content, order, timing, priority or importance of multiple objects, etc. For example, the first information and the second information do not represent differences in the amount of information, content, priority or importance, etc.
[0049] 4. The terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0050] 5. In each embodiment of the present application, "Network Element A sends Information A to Network Element B" can be understood as that the destination of Information A or the intermediate network element in the transmission path between the destination and the source is Network Element B, which may include sending Information A to Network Element B directly or indirectly. "Network Element B receives Information A from Network Element A" can be understood as that the source of Information A or the intermediate network element in the transmission path between the source and the destination is Network Element A, which may include receiving Information A from Network Element A directly or indirectly. Necessary processing may be performed on the information between the source and the destination of the information transmission, such as format change, etc., but the destination can understand the valid information from the source. Similar expressions in the present application can be understood similarly and will not be elaborated here.
[0051] In other words, the sending and receiving can be carried out between devices. For example, between Terminal Device #1 and Terminal Device #2, or can be carried out within a device. For example, sending or receiving between components, modules, chips, software modules or hardware modules within a device through a bus, wiring or interface.
[0052] 6. In the embodiments of the present application, the indication includes direct indication (also known as explicit indication) and implicit indication. Among them, directly indicating Information A means including Information A; implicitly indicating Information A means indicating Information A through the correspondence between Information A and Information B and directly indicating Information B. Among them, the correspondence between Information A and Information B can be predefined, pre-stored, pre-burned, or pre-configured.
[0053] 7. In the embodiments of the present application, Information C is used for the determination of Information D, which includes both the case where Information D is determined only based on Information C and the case where Information D is determined based on Information C and other information. In addition, when Information C is used for the determination of Information D, there may also be an indirect determination case. For example, Information D is determined based on Information E, and Information E is determined based on Information C.
[0054] 8. "Storage" or "saving" involved in the embodiments of the present application may refer to saving in one or more memories. The one or more memories may be set separately, or may be integrated in an encoder or decoder, a processor, or a communication device. The one or more memories may also be partly set separately and partly integrated in a decoder, a processor, or a communication device. The type of the memory may be any form of storage medium, which is not limited in the present application.
[0055] 9. The "protocol" involved in the embodiments of the present application may refer to a standard protocol in the communication field. For example, it may include the fourth-generation (4G) network / fifth-generation (5G) network protocol, the new radio (NR) protocol, and related protocols applied to future communication systems. The present application does not make any limitations in this regard.
[0056] 10. In the schematic diagrams in the attached drawings of the present application specification, the arrows or boxes shown by dotted lines represent optional steps or optional modules.
[0057] The technical solution provided by the present application can be applied to various communication systems, such as: the fifth-generation (5G) or new radio (NR) system, the long-term evolution (LTE) system, the LTE frequency division duplex (FDD) system, the LTE time division duplex (TDD) system, the wireless local area network (WLAN) system, the satellite communication system, future communication systems, such as the sixth-generation (6G) mobile communication system, or a fusion system of multiple systems, etc. The technical solution provided by the present application can also be applied to device-to-device (D2D) communication, vehicle-to-everything (V2X) communication, machine-to-machine (M2M) communication, machine type communication (MTC), and the Internet of Things (IoT) communication system or other communication systems.
[0058] A device in a communication system can send a signal to another device or receive a signal from another device. The signal can include information, signaling, data, etc. Herein, the device can also be replaced with an entity, a network entity, a communication device, a communication module, a node, a communication node, etc. In the present disclosure, the device is described as an example. For example, a communication system may include at least one terminal device and at least one network device. The network device can send a downlink signal to the terminal device, and / or the terminal device can send an uplink signal to the network device.
[0059] In the embodiments of the present application, the terminal device may also be referred to as a user equipment (UE), an access terminal, a user unit, a user station, a mobile station, a mobile device, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, a wireless communication device, a user agent, or a user device.
[0060] The terminal device may be a device that provides voice / data. For example, it may be a handheld device with wireless connection capabilities, a vehicle-mounted device, etc. Currently, some examples of terminals are: mobile phones, tablet computers, laptop computers, palmtop computers, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices, or other processing devices connected to a wireless modem, wearable devices, terminal devices in a 5G network, or terminal devices in a future evolved public land mobile network (PLMN). The embodiments of the present application are not limited thereto.
[0061] By way of example and not limitation, in the embodiments of the present application, the terminal device may also be a part of a network device for implementing the functions of a terminal device. For example, the network device may be an integrated access and backhaul (IAB) node. The IAB node integrates two parts: a mobile termination (MT) and a distributed unit (DU), or an MT and a base station (BS) part, where the BS includes a central unit (CU) and a DU. When the IAB node faces its parent node, it can be regarded as a terminal. At this time, the IAB node plays the role of the MT.
[0062] By way of example and not limitation, in the embodiments of the present application, the terminal device may also be a wearable device. A wearable device can also be called a wearable intelligent device, which is a general term for devices developed by applying wearable technologies to the intelligent design of daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is either directly worn on the body or integrated into the user's clothes or accessories. A wearable device is not just a hardware device, but also realizes powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable intelligent devices include those with complete functions and large sizes that can realize complete or partial functions without relying on a smart phone, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to cooperate with other devices such as smart phones, such as various smart bracelets and smart jewelry for physical sign monitoring.
[0063] In the embodiments of the present application, the device for implementing the functions of the terminal device may be the terminal device or a device capable of supporting the terminal device to implement such functions, such as a chip system. The device may be installed in the terminal device or used in matching with the terminal device. In the embodiments of the present application, the chip system may be composed of chips or may include chips and other discrete devices. In the embodiments of the present application, only the case where the device for implementing the functions of the terminal device is the terminal device is taken as an example for illustration, which does not limit the solutions of the embodiments of the present application.
[0064] The network device in the embodiments of this application can be a device used to communicate with terminal devices. This network device can also be referred to as an access network device or a radio access network device. For example, the network device can be a base station. The network device in the embodiments of this application can refer to a radio access network (RAN) node (or device) that connects terminal devices to a wireless network. A base station can generally cover various names as follows, or be replaced with the following names. For example: Node B, evolved Node B (eNB), next generation Node B (gNB), relay station, IAB node (such as the BS function part in an IAB node), access point, transmitting and receiving point (TRP), transmitting point (TP), master station, slave station, multi-mode radio (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), radio unit (RU), positioning node, etc. A base station can be a macro base station, a micro base station, a relay node, a donor node, or the like, or a combination thereof. A base station can also refer to a communication module, a modem, or a chip disposed in the foregoing device or apparatus. A base station can also be a mobile switching center and a device that undertakes the base station function in D2D, V2X, M2M communications, a network-side device in a 6G network, a device that undertakes the base station function in a future communication system, etc. A base station can support networks with the same or different access technologies. Optionally, the RAN node can also be a server, a wearable device, a vehicle, or a vehicle-mounted device, etc. For example, the access network device in vehicle to everything (V2X) technology can be a road side unit (RSU). The embodiments of this application do not limit the specific technologies and specific device forms adopted by the network device.
[0065] In some deployments, the network device mentioned in the embodiments of the present application may be a device including a CU, or a DU, or a device including a CU and a DU, or a control plane CU node (central unit-control plane, CU-CP) and a user plane CU node (central unit-user plane, CU-UP) and a DU node. For example, the network device may include a gNB-CU-CP, a gNB-CU-UP, and a gNB-DU.
[0066] In some deployments, multiple RAN nodes cooperate to assist a terminal in achieving wireless access, and different RAN nodes respectively implement some functions of a base station. For example, the RAN node may be a CU, a DU, a CU-CP, a CU-UP, or an RU, etc. The CU and the DU may be separately provided, or may also be included in the same network element, such as a BBU. The RU may be included in a radio frequency device or a radio frequency unit, such as included in an RRU, an AAU, or an RRH.
[0067] The RAN node may support one or more types of fronthaul interfaces. Different fronthaul interfaces respectively correspond to DUs and RUs with different functions. If the fronthaul interface between the DU and the RU is a common public radio interface (CPRI), the DU is configured to implement one or more of the baseband functions, and the RU is configured to implement one or more of the radio frequency functions. If the fronthaul interface between the DU and the RU is another interface, compared with the CPRI, some of the downlink and / or uplink baseband functions, for example, for the downlink, one or more of precoding, digital beamforming (BF), or inverse fast Fourier transform (IFFT) / adding cyclic prefix (CP), are moved from the DU to the RU for implementation. For the uplink, one or more of digital beamforming (BF), or fast Fourier transform (FFT) / removing cyclic prefix (CP) are moved from the DU to the RU for implementation. In a possible implementation manner, this interface may be an enhanced common public radio interface (eCPRI). In the eCPRI architecture, the splitting method between the DU and the RU is different, corresponding to different categories (Cat) of eCPRI, such as eCPRI Cat A, B, C, D, E, F.
[0068] Taking eCPRI Cat A as an example, for downlink transmission, with layer mapping as the division, the DU is configured to implement layer mapping and one or more functions before it (i.e., one or more of encoding, rate matching, scrambling, modulation, layer mapping), while other functions after layer mapping (such as one or more of RE mapping, digital beamforming (BF), or inverse fast Fourier transform (IFFT) / adding cyclic prefix (CP)) are moved to the RU for implementation. For uplink transmission, with de-RE mapping as the division, the DU is configured to implement demapping and one or more functions before it (i.e., one or more of decoding, derate matching, descrambling, demodulation, inverse discrete Fourier transform (IDFT), channel equalization, de-RE mapping), while other functions after demapping (such as one or more of digital BF or fast Fourier transform (FFT) / removing CP) are moved to the RU for implementation. It can be understood that for the function descriptions of the DU and RU corresponding to various types of eCPRI, reference can be made to the eCPRI protocol and will not be elaborated here.
[0069] In a possible design, the processing unit in the BBU for implementing baseband functions is called the baseband high (BBH) unit, and the processing unit in the RRU / AAU / RRH for implementing baseband functions is called the baseband low (BBL) unit.
[0070] In different systems, the CU (or CU-CP and CU-UP), DU, or RU may also have different names, but those skilled in the art can understand their meanings. For example, in the ORAN system, the CU can also be called O-CU (open CU), the DU can also be called O-DU, the CU-CP can also be called O-CU-CP, the CU-UP can also be called O-CU-UP, and the RU can also be called O-RU. Any one of the CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.
[0071] In the embodiments of the present application, the device for implementing the functions of a network device may be the network device itself; or it may be a device capable of supporting the network device to implement such functions, such as a chip system, a hardware circuit, a software module, or a combination of a hardware circuit and a software module. This device may be installed in the network device or used in conjunction with the network device. In the embodiments of the present application, only the case where the device for implementing the functions of the network device is the network device itself is taken as an example for illustration, which does not limit the solutions of the embodiments of the present application.
[0072] The network device and / or the terminal device may be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; they may also be deployed on water; or they may be deployed on airplanes, balloons, and satellites in the air. In the embodiments of the present application, the scenarios where the network device and the terminal device are located are not limited. In addition, the terminal device and the network device may be hardware devices, or software functions running on dedicated hardware or general hardware. For example, they may be virtualized functions instantiated on a platform (such as a cloud platform), or entities including dedicated or general hardware devices and software functions. The present application does not limit the specific forms of the terminal device and the network device.
[0073] In a wireless communication network, such as a mobile communication network, the services supported by the network are becoming increasingly diverse, and thus the requirements to be met are also becoming increasingly diverse. For example, the network needs to be able to support ultra-high speeds, ultra-low latency, and / or massive connections. This characteristic makes network planning, network configuration, and / or resource scheduling increasingly complex. In addition, due to the increasing power of the network, such as supporting higher and higher frequencies, supporting high-order multiple input multiple output (MIMO) technology, supporting beamforming, and / or supporting new technologies such as beam management, network energy saving has become a hot research topic. These new requirements, new scenarios, and new characteristics have brought unprecedented challenges to network planning, operation and maintenance, and efficient operation. To meet this challenge, artificial intelligence technology can be introduced into the wireless communication network to achieve network intelligence.
[0074] To support AI technology in the wireless network, AI nodes may also be introduced into the network.
[0075] Optionally, the AI node may be deployed at one or more of the following positions in the communication system: access network device, terminal device, or core network device, etc. Or, the AI node may also be deployed separately. For example, it may be deployed at a position outside any of the above-mentioned devices, such as the host of an over the top (OTT) system or a cloud server. The AI node can communicate with other devices in the communication system, and the other devices may be, for example, one or more of the following: network device, terminal device, or network element of the core network.
[0076] It can be understood that the number of AI nodes in the embodiments of the present application is not limited. For example, when there are multiple AI nodes, the multiple AI nodes can be divided based on functions, such as different AI nodes being responsible for different functions.
[0077] It can also be understood that the AI nodes can be independent devices, or can be integrated into the same device to implement different functions, can also be network elements in hardware devices, can also be software functions running on dedicated hardware, or virtualized functions instantiated on a platform (such as a cloud platform). The present application does not limit the specific form of the AI nodes. Among them, the AI nodes can be AI network elements or AI modules.
[0078] Figure 1 is a schematic diagram of an application framework 100 in a communication system applicable to the embodiments of the present application. As Figure 1 shown, the network elements in the communication system are connected through interfaces (such as NG, Xn), or the air interface. One or more AI modules are provided in one or more of these network element nodes, such as cloud devices, access network nodes (RAN nodes), and terminals (for clarity, Figure 1 only 1 is shown). The access network node can be a separate RAN node, or can include multiple RAN nodes. For example, it includes a CU and a DU. One or more AI modules can also be provided in the CU and / or the DU. Optionally, the CU can also be split into a CU-CP and a CU-UP. One or more AI models are provided in the CU-CP and / or the CU-UP.
[0079] The AI module is used to implement the corresponding AI function. The AI modules deployed in different network elements can be the same or different. The models of the AI module can implement different functions according to different parameter configurations. The models of the AI module can be configured based on one or more of the following parameters: structural parameters (such as at least one of the number of neural network layers, the width of the neural network, the connection relationship between layers, the weights of neurons, the activation function of neurons, or the bias in the activation function), input parameters (such as the type and / or dimension of the input parameters), or output parameters (such as the type and / or dimension of the output parameters). Among them, the bias in the activation function can also be referred to as the bias of the neural network. An AI module can have one or more models. One model can infer an output, and the output includes one parameter or multiple parameters. The learning process, training process, or inference process of different models can be deployed in different nodes or devices, or can be deployed in the same node or device.
[0080] In this communication system, the AI model of the access network node interacts with the AI model in the cloud to exchange necessary data and finally feedback data to the user. Among them, the data can be carried on a physical channel. For example, it can be carried on a physical downlink control channel (PDCCH), a physical downlink shared channel (PDSCH), a physical uplink shared channel (PUSCH), or a physical uplink control channel (PUCCH), etc. For another example, it can be carried on a physical sidelink control channel (PSCCH) or a physical sidelink shared channel (PSSCH).
[0081] In this communication system, the AI module in the cloud can be located in a cloud device or can be separated from the cloud device. Similarly, for the AI model of the access network device, the AI module is located in a network device or can be separated from the network device. This application does not limit this.
[0082] Figure 2 It is a schematic diagram of an application framework 200 in the communication system applicable to the embodiments of this application. As Figure 2 shown, the communication system includes a RAN intelligent controller (RIC). For example, the RIC is used to implement AI-related functions. The RIC includes a near-real time RIC (near-RT RIC) and a non-real time RIC (Non-RT RIC). Among them, the non-real time RIC mainly processes non-real time information, such as data that is not sensitive to latency, and the latency of this data can be in seconds. The real time RIC mainly processes near-real time information, such as data that is relatively sensitive to latency, and the latency of this data is in tens of milliseconds.
[0083] Near-real-time RIC is used for model training and inference. For example, it is used to train an AI model and perform inference using this AI model. The near-real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (such as CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or inference data. Optionally, the near-real-time RIC can deliver the inference result to the RAN node and / or the terminal. Optionally, the inference result can be exchanged between the CU and the DU, and / or between the DU and the RU. For example, the near-real-time RIC delivers the inference result to the DU, and the DU sends it to the RU.
[0084] Non-real-time RIC is also used for model training and inference. For example, it is used to train an AI model and perform inference using this model. The non-real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (such as CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or inference data, and the inference result can be delivered to the RAN node and / or the terminal. Optionally, the inference result can be exchanged between the CU and the DU, and / or between the DU and the RU. For example, the non-real-time RIC delivers the inference result to the DU, and the DU sends it to the RU.
[0085] The near-real-time RIC and the non-real-time RIC can also be separately set as a network element. Optionally, the near-real-time RIC and the non-real-time RIC can also be part of other devices. For example, the near-real-time RIC is set in the RAN node (such as in the CU or DU), while the non-real-time RIC is set in the OAM, cloud server, core network device, or other network devices.
[0086] Figure 3 It is a schematic diagram of the applicable communication system 300 in the embodiments of the present application. As Figure 3 shown, the communication system 300 may include at least one network device, for example, the network device 110; the communication system 100 may also include at least one terminal device, for example, the terminal device 120 and the terminal device 130. The network device 110 can communicate with the terminal devices (such as the terminal device 120 and the terminal device 130) through a wireless link. Between the communication devices in this communication system, for example, between the network device 110 and the terminal device 120, communication can be carried out through multi-antenna technology. The communication system 300 also includes an AI network element 140. The AI network element 140 is used to perform AI-related operations, for example, constructing a training data set or training an AI model, etc.
[0087] In a possible implementation, the network device 110 may send data related to the training of the AI model to the AI network element 140. The AI network element 140 constructs a training data set and trains the AI model. For example, the data related to the training of the AI model may include the data reported by the terminal device. The AI network element 140 may send the operation results related to the AI model to the network device 110 and forward them to the terminal device through the network device 110. For example, the operation results related to the AI model may include at least one of the following: the AI model that has completed training, the evaluation result of the model, or the test result, etc.
[0088] It should be understood that Figure 4 Taking the example that the AI network element 140 is directly connected to the network device 110 only for illustration, in other scenarios, the AI network element 140 may be connected to both the network device 110 and the terminal device simultaneously. Or, the AI network element 140 may also be connected to the network device 110 through a third-party network element. The embodiments of the present application do not limit the connection relationship between the AI network element and other network elements. For example, the AI network element 140 may also be set as a module in the network device, for example, set in Figure 3 the network device 110 shown.
[0089] It should be noted that Figure 1 and Figure 3 are simplified schematic diagrams shown only for easy understanding. For example, the communication system may further include other devices, such as wireless relay devices and / or wireless backhaul devices, etc. Figure 1 and Figure 3 are not drawn. In actual applications, Figure 1 and Figure 3 the communication system shown may include multiple network devices or multiple terminal devices. The embodiments of the present application do not limit the number of network devices and terminal devices included in the communication system.
[0090] It should also be noted that the system architecture and service scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of the network architecture and the emergence of new service scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0091] For easy understanding of the present application, some technical concepts involved in the present application are briefly described below.
[0092] (1) Machine learning (ML) is an important technical approach to realizing AI. ML can be divided into supervised learning, unsupervised learning, and reinforcement learning.
[0093] The deep neural network (DNN) is a specific implementation form of ML. According to the universal approximation theorem, theoretically, a neural network can approximate any continuous function, enabling the neural network to have the ability to learn any mapping. Traditional communication systems need to rely on rich expert knowledge to design communication modules, while deep learning communication systems based on DNN can automatically discover implicit pattern structures from large datasets, establish mapping relationships between data, and obtain performance superior to traditional modeling methods. The idea of DNN comes from the neuron structure of the brain tissue. Each neuron performs a weighted summation operation on its input values, and the weighted summation result passes through a non - linear function to generate an output. For the description of the neuron structure, please refer to Figure 4 .
[0094] As Figure 4 shown, Figure 4 is a schematic diagram of a neuron structure. As Figure 4 shown, assume the input of the neuron is x = [x 0 , …, x n , the weights corresponding to the input are w = [w 0 , …, w n , the bias of the weighted summation is b, and the form of the non - linear function can be diversified. An example is the max{0, x} maximum value function. Then the execution effect of a neuron can be
[0095] (2) AI model: The AI model is an algorithm or computer program that can implement AI functions. The AI model represents the mapping relationship between the input and output of the model. The types of AI models can be neural networks, linear regression models, decision tree models, support vector machines (SVMs), Bayesian networks, Q - learning models, or other ML models. The AI model in the embodiments of this application can be a dual - end model or a single - end model.
[0096] (3) Model inference: Using the trained model to solve practical problems.
[0097] (4) Natural language processing: Natural language processing is an interdisciplinary sub - field of computer science and linguistics. It mainly involves endowing computers with the ability to support and manipulate language and writing. Its goal is a computer that can "understand" the content of a document, including the subtle nuances of the language in the document. Then, this technology can accurately extract the information and insights contained in the document and classify and organize the document itself.
[0098] (5) Prompt: It is a natural language text that describes the tasks that AI should perform.
[0099] (6) Context: In linguistics, the context is the object or entity surrounding the focal event, and it is the framework that surrounds the event and provides resources for its proper interpretation. The meaning of words in natural language processing is usually inferred from the context in which they occur.
[0100] It can be understood that the implementation of the AI model can be a hardware circuit, software, or a combination of software and hardware, without limitation. Non-limiting examples of software include: program code, programs, subroutines, instructions, instruction sets, code, code segments, software modules, applications, or software applications, etc.
[0101] Recently, large language models (LLMs) have revolutionized the fields of natural language processing (NLP) and AI. An LLM is an AI model based on machine learning algorithms and natural language processing techniques that has pushed text generation, understanding, and interaction into a new era. An LLM is a language model characterized by its large size and capacity, and it is trained by preprocessing a large amount of text data collected from the Internet based on AI algorithms. It usually contains tens of millions to billions of weights and is designed to simulate and understand the characteristics and rules of human language. An LLM builds an extremely large-capacity neural network model to achieve multiple applications such as information retrieval, sentiment analysis, text generation, chatbots, and dialogue AI. It can be used to answer questions on any online platform, allowing visitors to receive meaningful information without contacting the expert corresponding to the question. For a user's question, an LLM-based intelligent agent is an AI-driven program that uses natural language processing and AI to infer the answer to the question. Different from most traditional question-and-answer applications, the performance of an LLM depends on its huge database, which contains texts from different sources, including books, blogs, and news articles. The extremely large-scale database enables it to understand information about various topics and provide better answers. It enables the machine to understand the prompt words provided by the user and offer useful help to the questioner. As the inference accuracy level of LLMs improves and the application scope expands, more and more users obtain answers to the problems they encounter and help with their daily tasks by interacting with this advanced application.
[0102] Since large language models (LLMs) have a huge number of parameters that require a large amount of storage space, and powerful hardware is needed to achieve fast computing during their inference process, current inference solutions for LLMs are often implemented based on cloud computing, that is, cloud-based LLMs inference. In this solution, an LLM application programming interface can be provided for users, and users (individuals and / or organizations) are allowed to access the LLM application from any device with an Internet connection and obtain inference results. The LLM application programming interface is essentially a connection between two devices, enabling them to have the ability to send information back and forth. The cloud provides the LLM to the user terminal, and the user interacts with the cloud through the user terminal to obtain AI intelligent services. The following briefly describes the interaction process between the user and the cloud, mainly including the following steps:
[0103] 1) The user sends a question prompt to the network device through the user terminal. For example, the question prompt sent by the user is: "Please recommend a library";
[0104] 2) The base station sends the user's question prompt to the cloud;
[0105] 3) The cloud runs the AI model to calculate the result corresponding to the user's request;
[0106] 4) The cloud feeds back the result to the base station; the base station downlinks the result fed back by the cloud to the user.
[0107] To reduce the latency of the AI model serving users, a cloud-network collaborative inference solution has been proposed currently. Specifically, by offloading the cloud-trained LLM to network devices (such as base stations), that is, offloading part of the inference tasks of the cloud AI model to network entity devices. This solution can prevent congestion caused by a large number of users accessing the cloud AI model to request inference services. In addition, this solution believes that network devices have sufficient computing power to perform AI model inference and are closer to users than cloud servers. Therefore, users can interact with the network device AI model, reducing the transmission latency to a certain extent.
[0108] Whether it is cloud-based LLM inference or cloud-network collaborative LLM inference, generally, users do not provide a comprehensive problem description. When the question prompt is relatively short, the inference accuracy is low. If precise services are to be achieved, the AI model needs to clearly locate the user's question. Therefore, the user needs to interact with the AI model multiple times, resulting in a large latency for the AI model to infer a satisfactory answer for the user, ultimately leading to a poor user experience.
[0109] In view of this, the present application provides a method for inference, which can effectively solve the above technical problems. The following combines Figure 5 to introduce the method proposed in the present application in detail.
[0110] Figure 5 It is a schematic flowchart of a communication method 500 provided by this application. The method 500 includes the following steps.
[0111] S510, the terminal device sends a first prompt word to the first AI model. The first prompt word is used to request a corresponding reply based on the first prompt word. The first AI model receives the first prompt word from the terminal device.
[0112] Exemplarily, if the user using the terminal device wants to search for a library, the first prompt word can be "Please recommend a library".
[0113] Exemplarily, if the user using the terminal device wants to search for a park with the destination of Beijing, the first prompt word can be "Please recommend a park in Beijing".
[0114] S520, the first AI model obtains a second prompt word based on the first prompt word and the local information of the first network device. The second prompt word is used to describe the information corresponding to at least one parameter related to the first prompt word in the local information of the first network device. The local information of the first network device is the characteristic information of the cells covered by the first network device. The first network device is the network device in the geographical range related to the first prompt word.
[0115] It can be understood that if the first prompt word does not include geographical range information (the first prompt word can be "Please recommend a library"), or the geographical range information included in the first prompt word does not meet the geographical range accuracy required for reasoning, the first AI model needs to interact with the user further until the geographical range accuracy required for reasoning is obtained. Among them, the geographical range accuracy required for reasoning can be understood as that the first AI model determines that the accuracy of the geographical range related to the first prompt word obtained is not less than the geographical range accuracy required for reasoning. For example, the geographical range accuracy required for reasoning is a city, and the geographical range accuracy that meets the requirements for reasoning is a certain city or a certain jurisdiction within a certain city. Another example is that the geographical range accuracy required for reasoning is a jurisdiction of a city, and the geographical range accuracy that meets the requirements for reasoning is a certain jurisdiction of a certain city, or a certain area within a certain jurisdiction of a certain city.
[0116] Optionally, in the above scenario, the method further includes: the first AI model sends a fifth prompt word to the terminal device, and the fifth prompt word is used to interact with the user to further obtain geographical range information that meets the geographical range accuracy. Correspondingly, the terminal device receives the fifth prompt word from the first AI model. After that, the user sends a seventh prompt word to the first AI model through the terminal device, and the seventh prompt word is used to describe the geographical range information related to the first prompt word. Correspondingly, the first AI model receives the seventh prompt word from the terminal device. After that, the first AI model obtains the local information of the network device (i.e., the first network device) corresponding to the corresponding geographical range based on the seventh prompt word.
[0117] Exemplarily, the fifth prompt word can be used to prompt the input of geographical range information related to the first prompt word. For example, if the first prompt word is "Please recommend a library", the fifth prompt word can be "Enter the approximate range of the library you want to recommend, for example, Pudong New Area, Shanghai, Haidian District, Beijing, etc."
[0118] Exemplarily, the fifth prompt word can also include a sixth prompt word and multiple alternative geographical range information. The sixth prompt word is used to prompt to select the geographical range information related to the first prompt word from the multiple alternative geographical range information. For example, if the first prompt word is "Please recommend a library", the fifth prompt word can be "Please select the approximate range of the library you want to recommend from the following options", and the sixth prompt word can include options corresponding to multiple geographical ranges, and the user can select the option corresponding to the geographical range they expect related to the first prompt word.
[0119] Then, when the first AI model obtains the geographical range information that meets the required accuracy for reasoning, it can obtain the second prompt word based on the local information of the network device (i.e., the first network device) corresponding to the obtained geographical range information, where the second prompt word is used to describe the information corresponding to at least one parameter related to the first prompt word in the local information of the first network device. The following are two implementation methods for obtaining the second prompt word based on the first prompt word and the local information of the first network device.
[0120] Implementation method 1, the second prompt word can be regarded as the first AI model performing information mining on the basis of the first prompt word (mining at least one parameter related to the first prompt word), and supplementing the mined information based on the local information of the first network device (obtaining information related to the at least one mined parameter from the local data of the first network device), so as to improve the first prompt word for better derivation and reply. The following is an example.
[0121] For example, the first prompt word can be "Please recommend a library", and the parameter related to the first prompt word can be wireless quality data. Suppose after the first AI model interacts with the user, it is clear that the user needs to recommend a library in Pudong New Area where the current user is located. Then the first AI model obtains the information related to the wireless quality data stored in the local information of the network devices (i.e., the first network devices) within the scope of Pudong New Area. The information related to the wireless quality data can include, but is not limited to, the following parameters: wireless traffic data, the number of active users, and user CQI. For example, the second prompt word can be "The current base station is located in Pudong New Area. The wireless traffic on Monday in this area is a1, the number of access users is b1, the wireless traffic on Tuesday is a2, the number of access users is b2,..., the future wireless traffic is x, and the number of access users is u".
[0122] Implementation method 2: The first AI model infers multiple reply results based on the first prompt word. Then, the first AI model conducts information mining on the basis of the first prompt word (mining at least one parameter related to the first prompt word). After that, it supplements the mined information for the multiple reply results based on the local information of the first network device (obtaining the information related to at least one parameter related to each reply result from the local data of the first network device), so as to improve the multiple reply results in order to determine the optimal reply result. The following is an example for illustration.
[0123] For example, the first prompt word can be "Please recommend a library". Suppose after the first AI model interacts with the user, it is clear that the user needs to recommend a library in Pudong New Area where the current user is located. The first AI infers multiple libraries in Pudong New Area (i.e., multiple reply results). For example, the multiple libraries in Pudong New Area include "Library A, Library B, and Library C". Then, the first AI model obtains the information related to the wireless quality data related to the multiple reply results stored in the local information of the network devices (i.e., the first network devices) within the scope of Pudong New Area. For example, the second prompt word can be "The current base station is located in Pudong New Area. The wireless traffic on Monday in the area where Library A is located is a11, the number of access users is b11, the wireless traffic on Tuesday is a12, the number of access users is b12,..., the future wireless traffic is x1, the number of access users is u1, the wireless traffic on Monday in the area where Library B is located is a21, the number of access users is b21, the wireless traffic on Tuesday is a22, the number of access users is b22,..., the future wireless traffic is x2, the number of access users is u2, the wireless traffic on Monday in the area where Library C is located is a31, the number of access users is b31, the wireless traffic on Tuesday is a32, the number of access users is b32,..., the future wireless traffic is x3, the number of access users is u3.
[0124] Optionally, before S520, the method further includes: the first AI model obtains a fourth prompt word based on the first prompt word, and the fourth prompt word is used to describe at least one parameter related to the first prompt word. For example, if the first prompt word is "Please recommend a library", the fourth prompt word can be "Also consider the wireless service quality". For example, if the first prompt word can be "Please recommend a park", the fourth prompt word can be "Also consider the environmental situation". Exemplarily, the information related to environmental perception data in the local information of the first network device may include, but is not limited to, information such as temperature, humidity, air pressure, and wind speed.
[0125] After that, the first AI model obtains a second prompt word based on the first prompt word, the local information of the first network device, and the fourth prompt word.
[0126] It should be understood that this application does not limit whether the first AI model generates the fourth prompt word. The first AI model can also skip generating the fourth prompt word and directly obtain the second prompt word, and this application does not limit this.
[0127] It can be understood that in Implementation Mode 1, the local information of the first network device is used to supplement the description of the first prompt word to improve the context information of the first prompt word, so that the subsequent AI model for reasoning can better reason based on the prompt word after supplementing the context information and obtain a more accurate reply result. In the second implementation mode, based on the first prompt word, the first AI model first roughly derives multiple reply results for the first prompt word. After that, the local information of the first network device is used to supplement the description of the multiple reply results to improve the context information of the corresponding reply, so that the subsequent AI model for reasoning can select the best reply result from the multiple results based on the multiple reply results after supplementing the context information.
[0128] Exemplarily, the subsequent AI model for reasoning about the reply result can be the first AI model or other AI models.
[0129] It should be understood that if the subsequent AI model for reasoning about the reply result is the first AI model, then S530 is executed. Exemplarily, the first AI model can be the AI model of the network device of the service terminal device, or the second AI model can be the AI model in the cloud.
[0130] It should be understood that if the subsequent AI model for reasoning about the reply result is the second AI model, then S540 is executed. Exemplarily, the first AI model can be the AI model of the network device in the cloud network collaborative reasoning scenario, and the second AI model is the AI model in the cloud in the cloud network collaborative reasoning scenario.
[0131] Exemplarily, the first AI model and the second AI model are LLMs.
[0132] S530, the first AI model sends the first reply result to the terminal device, where the first reply result is determined based on the third prompt word, and the third prompt word includes the first prompt word and the second prompt word. Correspondingly, the first AI model receives the first reply result from the terminal device.
[0133] Exemplarily, the second prompt word may further include a fourth prompt word.
[0134] Based on the above implementation method 1, exemplarily, if the first prompt word is "Please recommend a library", the third prompt word may be "Please recommend a library, taking into account the wireless quality data as well. The current base station is located in Pudong New Area, where the wireless traffic on Monday is a1, the number of access users is b1, the wireless traffic on Tuesday is a2, the number of access users is b2,..., the future wireless traffic is x, and the number of access users is u". Then, the first AI model can perform reasoning based on the third prompt word after supplementing the context information, and send the inferred first reply result to the terminal device (user).
[0135] Based on the above implementation method 2, exemplarily, if the first prompt word is "Please recommend a library", the third prompt word may be "Please recommend a library, taking into account the wireless quality data as well. The current base station is located in Pudong New Area. For the area where Library A is located, the wireless traffic on Monday is a11, the number of access users is b11, the wireless traffic on Tuesday is a12, the number of access users is b12,..., the future wireless traffic is x1, and the number of access users is u1. For the area where Library B is located, the wireless traffic on Monday is a21, the number of access users is b21, the wireless traffic on Tuesday is a22, the number of access users is b22,..., the future wireless traffic is x2, and the number of access users is u2. For the area where Library C is located, the wireless traffic on Monday is a31, the number of access users is b31, the wireless traffic on Tuesday is a32, the number of access users is b32,..., the future wireless traffic is x3, and the number of access users is u3". Then, the first AI model can perform reasoning based on multiple reply results after supplementing the context information, and send one best reply result (the first reply result) selected from the multiple reply results to the terminal device (user).
[0136] Optionally, since multiple replies corresponding to the first prompt word have been given in implementation method 2, the third prompt word may not include the first prompt word, and this application does not limit this.
[0137] In a possible implementation method, the first AI model can use the local information of the obtained first network device to verify its reasoning result, ensuring that the reply finally sent to the user is a reasonable reply, and avoiding reducing the user experience due to the first AI model obtaining an inappropriate reply result due to hallucination. For example, the user hopes to recommend a library in Pudong New Area of Shanghai, but the reply result finally given by the first AI model is a library in Jing'an District of Shanghai.
[0138] Optionally, before S530, the method further includes: the first AI model determines that the first reply result is an appropriate reply result based on the verification of the local information of the first network device.
[0139] Optionally, before S530, the method further includes: the first AI model infers a fourth reply result based on the third prompt word; the first AI model verifies that the fourth reply result is an inappropriate reply result based on the local information of the first network device. After that, the first AI model can re-infer a reply result based on the third prompt word until the first AI verifies that the reply result is an appropriate reply result.
[0140] S540, the first AI model sends the third prompt word to the second AI model. Correspondingly, the second AI model receives the third prompt word from the first AI model.
[0141] For the description of the third prompt word, see the description in S530, which will not be repeated here.
[0142] Optionally, after the second AI model obtains the third prompt word, the method further includes:
[0143] S550, the second AI model sends Reply Result #1 to the first AI model, where Reply Result #1 is the reply result inferred by the second AI model based on the third prompt word. Correspondingly, the first AI model receives Reply Result #1 from the second AI model.
[0144] S560, the first AI model sends Reply Result #1 to the terminal device.
[0145] Optionally, in this scenario, the first AI model can also use the obtained local information of the first network device to verify the inference result of the second AI model, ensuring that the reply finally sent to the user is a reasonable reply, and avoiding reducing the user experience due to the second AI model obtaining an inappropriate reply result due to the hallucination phenomenon. In this implementation, the second AI model is allowed to interact with the first AI model. Compared with the second AI model that cannot obtain the unique local information of the first network device for independent verification, the collaborative verification of the inference reply result can better ensure the effectiveness of the result fed back to the user.
[0146] Then, based on the collaborative verification method proposed above, after the first AI model receives the reply result #1, the method further includes: the first AI model verifies that the reply result #1 is an appropriate reply result based on the local information of the first network device, and the first AI model sends the reply result #1 to the terminal device. Correspondingly, the terminal device receives the reply result #1 from the first AI model; or, if the first AI model verifies that the reply result #1 is an inappropriate reply result based on the local information of the first network device, the first AI model does not execute S560, and the first AI model sends an eighth prompt word to the second AI model. The eighth prompt word includes a third prompt word and a ninth prompt word, and the ninth prompt word is used to prompt to re-infer the reply result based on the third prompt word. Correspondingly, the second AI model receives the eighth prompt word from the first AI model. After that, the second AI model re-infers a reply result #2 based on the third prompt word and sends it to the first AI model, and the first AI model continues to verify whether the reply result #2 is an appropriate reply result based on the local information of the first network device.
[0147] Based on the above technical solution, the first AI model can use the local information of the first network device related to the first prompt word for accurate inference, or assist the second AI model in accurate inference, reducing the number of times of asking the user questions. In addition, verifying the inference result based on the local data of the first network device can further ensure the user's demand for service quality on the basis of accurate inference, improving the user experience.
[0148] Next, the method 500 will be further described. In the method 500, the first AI model can store the first sample corresponding to the final reply result of the first prompt word sent to the user terminal in the local sample library, where the first sample is composed of <"the first prompt word", "the third prompt word", "the final reply result of the first prompt word">.
[0149] It can be understood that the local sample library of the first AI model also stores samples corresponding to multiple prompt words previously input by the user terminal. Then, for method 500, in S510, when the first AI model receives the first prompt word from the user terminal, it can select multiple samples related to the first prompt word from the historical samples of the user in the sample library and add them before the first prompt word. After that, in S520, the first AI model obtains the second prompt word based on the first prompt word, the local information of the first network device, and the tenth prompt word, where the tenth prompt word is used to describe at least one historical sample related to the first prompt word among the historical samples of the user included in the sample library of the first AI model. In addition, the third prompt word used for the first AI model to deduce the reply result in S530 and the third prompt word used for the second AI model to deduce the reply result in S540 include the tenth prompt word, the first prompt word, and the second prompt word. The subsequent steps are the same and will not be elaborated here one by one. This solution enriches the context information of the user's prompt word based on the user's historical samples, enabling the AI model for inferring the reply result to also provide accurate personalized services for the user based on information related to the user's habits or preferences during inference, thereby further enhancing the user experience.
[0150] To facilitate the understanding of the solution shown in method 500, the following takes the scenario of cloud-network collaborative inference as an example to describe method 500 based on implementation method 2 in S520.
[0151] Figure 6 It is a schematic diagram of a cloud-network collaborative inference method 600 proposed in this application. It should be understood that method 600 is only an implementation process of a possible specific scenario of cloud-network collaborative derivation. Among them, the AI model of the network device can be regarded as the first AI model in method 500, and the AI model in the cloud can be regarded as the second AI model in method 500. This method includes the following steps.
[0152] S601, the user terminal sends prompt word #1 (i.e., an example of the first prompt word) to the AI model of the network device, and prompt word #1 is "Please recommend a library". Correspondingly, the AI model of the network device receives prompt word #1 from the user terminal.
[0153] S602, the AI model of the network device determines that prompt word #1 does not contain geographical range information and sends prompt word #2 (i.e., an example of the fifth prompt word) to the user terminal. Prompt word #2 is used to further interact with the user to obtain geographical range information that meets the geographical range accuracy.
[0154] 603, the user terminal sends prompt word #3 (i.e., the seventh prompt word) to the AI model of the network device. Prompt word #3 is used to describe the geographical range information related to prompt word #1. Correspondingly, the AI model of the network device receives prompt word #3 from the terminal device.
[0155] For ease of description, here, taking the geographical range information indicated by the prompt #3 as an example to describe the area #1, and the area #1 meets the geographical range accuracy required for the inference response result.
[0156] S604, the AI model of the network device infers the prompt #4 (i.e., an example of the fourth prompt) based on the prompt #1.
[0157] Exemplarily, if the first AI model locally stores the historical examples of the user, the first AI model can select at least one example related to the prompt #1 from the historical examples of the user and add it before the prompt #1. If historical examples are added, the prompt #1 in the following text can be considered to include the selected historical examples and the prompt #1.
[0158] The AI model of the network device takes the received prompt #1 as the input of the AI model to dig deeper into the user's needs behind the prompt. Exemplarily, the AI model of the network device infers from the "library" in the prompt that the user will mainly be active indoors and may use wireless networks to perform operations such as online access and retrieval of books, in-library navigation, and electronic cloud notes, so the prompt #4 "also consider the wireless service quality" is derived and generated.
[0159] Optionally, the AI model of the network device can be a long short term memory (LSTM) model with a small capacity, or a Generative Pre-Trained Transformer (GPT), a model with a larger capacity such as bidirectional encoder representations from transformers (BERT) from the transformation module. As long as the storage space and computing power of the network device permit, the present application does not limit the specific size and scale of the AI model.
[0160] Figure 7 It is a schematic structural diagram of a possible AI model of the network device. This model is a GPT model. Optionally, the AI model parameters of the network device are generated by the network device through a certain strategy. For example, they can be generated through pre-training, downloaded from the cloud, or obtained from other third-party entities. As Figure 7As shown, the AI model includes 12 transformation modules (TransformerBlock), and each Transformer Block includes a masked self-attention module, a layer normalization module, a feed-forward neural network module, and a layer normalization module. The user prompt is used as the input of the first Transformer Block for learning. After that, the output of the first Transformer Block is used as the input of the second Transformer Block for continuous learning until the last Transformer Block infers an output.
[0161] S605. The AI model of the network device infers Prompt #5 (i.e., an example of the second prompt) based on Prompt #4 and the local data #1 of Network Device #1 (i.e., an example of the first network device), where Network Device #1 is the network device corresponding to Region #1.
[0162] Exemplarily, Prompt #4 is "Also consider the wireless service quality", and by invoking the specific data related to wireless service data in Data #1, the inferred Prompt #5 can be "The current base station is located in Region #1, where the wireless traffic on Monday is a1, the number of access users is b1, the wireless traffic on Tuesday is a2, the number of access users is b2,..., the future wireless traffic is x, and the number of access users is u".
[0163] Optionally, before the AI model of the network device invokes the data #1 of Network Device #1, the AI model of the network device also needs to request the user to authorize it to use the local data of the corresponding network device during inference. After the user completes the authorization, the network device can invoke the local database of the corresponding network device (such as Network Device #1) when using the AI model of the network device to perform inference.
[0164] S606. The AI model of the network device sends Prompt #6 (i.e., an example of the third prompt) to the AI model in the cloud. Correspondingly, the AI model in the cloud receives Prompt #6 from the AI model of the network device.
[0165] Exemplarily, Prompt #6 includes Prompt #1, Prompt #4, and Prompt #5, that is, Prompt #6 is "Please recommend a library. Also consider the wireless service quality. The current base station is located in Region #1, where the wireless traffic on Monday is a1, the number of access users is b1, the wireless traffic on Tuesday is a2, the number of access users is b2,..., the future wireless traffic is x, and the number of access users is u". It can be seen that the AI model in the cloud makes a supplementary description for Prompt #1, and supplements the context information of Prompt #1 in the finally sent Prompt #6 to the AI model in the cloud. This method can help the AI model in the cloud infer a more accurate response result.
[0166] S607, the AI model in the cloud infers the reply result R1 according to the prompt #6.
[0167] S608, the AI model in the cloud sends the reply result R1 to the AI model of the network device. Correspondingly, the AI model of the network device receives the reply result R1 from the AI model in the cloud.
[0168] S609, the AI model of the network device verifies based on Data #1 that the reply result R1 is an inappropriate reply result. Continue to execute S610.
[0169] S610, the AI model of the network device sends the prompt #7 (i.e., an example of the eighth prompt) to the AI model in the cloud. The prompt #7 includes the prompt #6 and the prompt #8 (i.e., an example of the ninth prompt). The prompt #8 is used to prompt to re-infer the reply result based on the prompt #6. For example, the prompt #8 is "Please provide a reply result other than R1". Correspondingly, the AI model in the cloud receives the prompt #7 from the AI model of the network device.
[0170] S611, the AI model in the cloud infers the reply result R2 according to the prompt #6.
[0171] S612, the AI model in the cloud sends the reply result R2 to the AI model of the network device. Correspondingly, the AI model of the network device receives the reply result R2 from the AI model in the cloud.
[0172] S613, the AI model of the network device verifies based on Data #1 that the reply result R2 is an appropriate reply result.
[0173] S614, the AI model of the network device sends the reply result R2 to the user terminal. Correspondingly, the user terminal receives the reply result R2 from the AI model of the network device.
[0174] It can be understood that method 600 is described in detail with the prompt word #1 being "recommend a library" as an example. For example, if the prompt word #1 is "please recommend a park", then based on this prompt word, the difference from method 600 is only that the context information to be supplemented in the derivation process is different. Specifically, in S604, the AI model of the network device infers from the "park" in the prompt word #1 that the user will mainly be outdoors, so the prompt word #4 "also consider the environmental conditions" is derived and generated. In S605, based on the prompt word #4 "also consider the environmental conditions", the AI model of the network device calls the specific data related to the environmental perception data in data #1, and the prompt word #5 inferred can be "the current base station is located in area #1, the temperature at 8 o'clock in this area is a1, the humidity is b1,... the temperature at 9 o'clock is a2, the humidity is b2,..., the future temperature at 8 o'clock in this area is x1, the humidity is x1,... the temperature at 9 o'clock is x2, the humidity is x2,...". In S606, the AI model of the network device sends the prompt word #6 to the AI model in the cloud, and the prompt word #6 is "please recommend a park. Also consider the environmental conditions, the current base station is located in area #1,, the temperature at 8 o'clock in this area is a1, the humidity is b1,... the temperature at 9 o'clock is a2, the humidity is b2,..., the future temperature at 8 o'clock in this area is x1, the humidity is y1,... the temperature at 9 o'clock is x2, the humidity is y2,...". The remaining steps are similar to method 600 and will not be elaborated here one by one.
[0175] It can be understood that some optional features in the embodiments of the present application can, in some scenarios, be independent of other features, and in some scenarios, can be combined with other features, without limitation.
[0176] It can also be understood that the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0177] It can also be understood that in the above some embodiments, the devices in the existing network architecture are mainly used as examples for illustrative purposes. It should be understood that the specific form of the devices is not limited in the embodiments of the present application. For example, devices that can achieve the same functions in the future are applicable to the embodiments of the present application.
[0178] It can also be understood that in the above method embodiments, the methods and operations implemented by the first AI model can also be implemented by components (such as chips or circuits) of an AI model.
[0179] Above, in combination with Figures 1 to 7The method provided by the embodiments of the present application is described in detail. The above method is mainly introduced from the perspective of the interaction between the first AI model and the first AI model and the second AI model. It can be understood that in order to implement the above functions, the first AI model includes the corresponding hardware structure and / or software module for executing each function.
[0180] Those skilled in the art should be able to realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0181] Hereinafter, in combination with Figure 8 and Figure 9 The communication device provided by the embodiments of the present application is described in detail. It should be understood that the description of the device embodiments corresponds to the description of the method embodiments. Therefore, for the content not described in detail, reference can be made to the above method embodiments. For the sake of brevity, some content will not be repeated. The embodiments of the present application can divide the functional modules of the first AI model according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation. Hereinafter, taking the division of each functional module corresponding to each function as an example for description.
[0182] The communication method provided by the present application has been described in detail above. Next, the communication device provided by the present application is introduced. In a possible implementation manner, the device is used to implement the steps or processes corresponding to the first AI model in the above method embodiments.
[0183] Figure 8 is a schematic block diagram of the communication device 800 provided by the embodiments of the present application. As Figure 8 shown, the device 800 may include a communication unit 810 and a processing unit 820. The communication unit 810 can communicate with the outside, and the processing unit 820 is used for data processing. The communication unit 810 can also be referred to as a communication interface or a transceiver unit.
[0184] In a possible design, the apparatus 800 may implement the steps or processes corresponding to those performed by the first AI model in the foregoing method embodiments. Among them, the processing unit 820 is configured to perform operations related to the processing of the first AI model in the foregoing method embodiments, and the communication unit 810 is configured to perform operations related to the transmission of the first AI model in the foregoing method embodiments.
[0185] It should be understood that the apparatus 800 here is embodied in the form of functional units. The term "unit" here may refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group of processors, etc.) for executing one or more software or firmware programs, and a memory, a combined logic circuit, and / or other suitable components that support the described functions. In an alternative example, those skilled in the art can understand that the apparatus 800 may specifically be the first AI model in the foregoing embodiments, and may be used to execute the respective processes and / or steps corresponding to the first AI model in the foregoing method embodiments. To avoid repetition, details are not described herein again.
[0186] The apparatus 800 in the foregoing various solutions has the function of implementing the corresponding steps performed by the first AI model in the foregoing method. The function may be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the foregoing functions; for example, the communication unit may be replaced by a transceiver (for example, the sending unit in the communication unit may be replaced by a transmitter, and the receiving unit in the communication unit may be replaced by a receiver), and other units, such as the processing unit, etc., may be replaced by a processor to respectively execute the transceiver operations and related processing operations in each method embodiment.
[0187] In addition, the foregoing communication unit may also be a transceiver circuit (for example, it may include a receiving circuit and a sending circuit), and the processing unit may be a processing circuit. In the embodiments of the present application, Figure 8 the apparatus may be the first AI model in the foregoing embodiments, or may be a chip or a chip system, for example: a system on chip (SoC). Among them, the communication unit may be an input / output circuit, a communication interface; the processing unit is a processor, a microprocessor, or an integrated circuit integrated on the chip. No limitation is made herein.
[0188] Figure 9Schematic block diagram of communication device 900 provided by an embodiment of the present application. The device 900 includes a processing circuit 910 and a transceiver circuit 920. Among them, the processing circuit 910 and the transceiver circuit 920 communicate with each other through an internal connection path. The processing circuit 910 is used to execute instructions to control the transceiver circuit 920 to send signals and / or receive signals. Among them, the transceiver circuit 920 can also be referred to as a transceiver interface.
[0189] Optionally, the device 900 may further include a memory 930. The memory 930 communicates with the processing circuit 910 and the transceiver circuit 920 through an internal connection path. The memory 930 is used to store instructions, and the processing circuit 910 can execute the instructions stored in the memory 930. In a possible implementation manner, the device 900 is used to implement each process and step corresponding to the first AI model in the above method embodiment.
[0190] It should be understood that the device 900 may specifically be the first AI model in the above embodiment, or a chip or a chip system. Correspondingly, the transceiver circuit 920 may be the transceiver interface of the chip, which is not limited herein. Specifically, the device 900 may be used to execute each step and / or process corresponding to the first AI model in the above method embodiment. Optionally, the memory 930 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type. The processing circuit 910 may be used to execute the instructions stored in the memory, and when the processing circuit 910 executes the instructions stored in the memory, the processing circuit 910 is used to execute each step and / or process of the above method embodiment corresponding to the first AI model.
[0191] In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0192] It should be noted that the processing circuit in the embodiments of the present application may be the processing circuit in an integrated circuit chip and has the ability to process signals. In the implementation process, each step of the above method embodiments can be completed by the integrated logic circuit of the hardware in the processing circuit or the instructions in the form of software. The above processing circuit may be the following devices or included in the following devices: general-purpose processor, digital signal processing (DSP), ASIC, field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor in the embodiments of the present application can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processing circuit reads the information in the memory and combines its hardware to complete the steps of the above method.
[0193] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory
[0194] It should be noted that when the processing circuit is a processor or is included in a processor, and the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate, or transistor logic device, discrete hardware component, the memory (storage module) can be integrated in the processor.
[0195] In addition, the present application also provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions are run on a computer, the operations and / or processes performed by the first AI model in the method embodiments of the present application are executed.
[0196] The present application also provides a computer program product, which includes computer program code or instructions. When the computer program code or instructions are run on a computer, the operations and / or processes performed by the first AI model in the method embodiments of the present application are executed.
[0197] In addition, the present application also provides a chip, which includes a processing circuit. A memory for storing a computer program is provided independently of the chip or disposed within the chip. The processing circuit is configured to execute the computer program stored in the memory, so that the operations and / or processes performed by the first AI model in any of the method embodiments are executed.
[0198] Further, the chip may further include a communication interface. The communication interface may be an input / output interface or an interface circuit, etc. Further, the chip may further include a memory.
[0199] In addition, the present application also provides a communication system, including the first AI model and a terminal device in the embodiments of the present application. Optionally, the communication system further includes a second AI model.
[0200] It should also be noted that the memories described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0201] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0202] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the essence of the technical solution of this application, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0203] It should be understood that the "embodiments" mentioned throughout the specification mean that specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the embodiments throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.
[0204] It should also be understood that in the present application, "when", "if" and "in case" all refer to the fact that the network element will perform corresponding processing under certain objective circumstances, which does not limit the time, and does not require the network element to have a judgment action when implemented, nor does it mean that there are other limitations.
[0205] It should also be understood that in the present application, "at least one" means one or more, and "a plurality" means two or more. "At least one item (or one)" or its similar expressions mean one item (or one) or more items (or one), that is, any combination of these items, including any combination of a single item (or one) or plural items (or one). For example, at least one item (or one) of a, b, or c means: a, b, c, a and b, a and c, b and c, or a, b and c.
[0206] It should also be understood that the term "and / or" in this document is merely an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, including A and B, and B exists alone, where A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. For example, A / B means: A or B.
[0207] It should also be understood that in each embodiment of the present application, "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
Claims
1. A method of reasoning, characterized in that, it includes: The first artificial intelligence (AI) model receives a first prompt word from a terminal device, and the first prompt word is used to request a corresponding reply based on the first prompt word; The first AI model obtains a second prompt word based on the first prompt word and the local information of the first network device. The second prompt word is used to describe information corresponding to at least one parameter related to the first prompt word in the local information of the first network device. The local information of the first network device is the characteristic information of the cell covered by the first network device, and the first network device is a network device within the geographical range related to the first prompt word; The first AI model sends a third prompt word to the second AI model. The second AI model is an AI model for reasoning the reply corresponding to the first prompt word, and the third prompt word includes the first prompt word and the second prompt word, or, The first AI model sends a first reply result to the terminal device based on the third prompt word.
2. The method according to claim 1, characterized in that, The first AI model obtains a fourth prompt word based on the first prompt word, and the fourth prompt word is used to describe at least one parameter related to the first prompt word.
3. The method according to claim 2, characterized in that, The third prompt word further includes the fourth prompt word.
4. The method according to any one of claims 1 to 3, characterized in that, If the first prompt word does not include geographical range information or the geographical range information included in the first prompt word does not meet the geographical range accuracy required for reasoning, before the first AI model obtains the second prompt word, the method further includes: The first AI model sends a fifth prompt word to the terminal device. The fifth prompt word is used to prompt the input of geographical range information related to the first prompt word, or the fifth prompt word includes a sixth prompt word and multiple alternative geographical range information, and the sixth prompt word is used to prompt to select geographical range information related to the first prompt word from the multiple alternative geographical range information; The first AI model receives a seventh prompt word from the terminal device, and the seventh prompt word is used to indicate geographical range information related to the first prompt word; The first AI model obtains the local information of the first network device based on the seventh prompt word.
5. The method according to any one of claims 1 to 4, characterized in that, When the first AI model sends the fourth prompt word to the second AI model, the method further includes: The first AI model receives a second reply result from the second AI model. The second reply result is the reply result inferred by the second model based on the third prompt word; The first AI model verifies that the second reply result is a reasonable reply result based on the local information of the first network device; The first AI model sends the second reply result to the terminal device.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: The first AI model receives a third reply result from the second AI model, and the third reply result is a reply result obtained by the second AI model based on the third prompt word. The first AI model verifies that the third reply result is an unreasonable reply result based on the local information of the first network device. The first AI model sends an eighth prompt word to the second AI model, and the eighth prompt word includes the third prompt word and a ninth prompt word, where the ninth prompt word is used to prompt to re-infer the reply result based on the third prompt word.
7. The method according to any one of claims 1 to 6, wherein, Before the first AI model sends a first result to the terminal device based on the third prompt word, the method further includes: The first AI model verifies that the first result is a reasonable reply result based on the local information of the first network device.
8. The method according to claim 7, wherein, Before the first AI model sends a first reply result to the terminal device based on the third prompt word, the method further includes: The first AI model infers a fourth reply result based on the third prompt word; The first AI model verifies that the fourth reply result is an unreasonable reply result based on the local information of the first network device.
9. The method according to any one of claims 1 to 8, wherein, The method further includes: The first AI model saves a first sample in the local sample library, and the first sample includes the third prompt word and the final reply result of the third prompt word.
10. The method according to claim 9, wherein, The first AI model obtains a second prompt word based on the first prompt word and the local information of the first network device, including: The first AI model obtains the second prompt word based on the first prompt word, the local information of the first network device, and a tenth prompt word, where the tenth prompt word is used to describe at least one sample in the sample library related to the first prompt word.
11. The method according to any one of claims 1 to 10, wherein, The first AI model is an AI model for inferring the reply corresponding to the first prompt word, and the first AI model is an AI model in the cloud or an AI model of a network device serving the terminal device, or, The second AI model is an AI model for inferring the reply corresponding to the first prompt word, the first AI model is an AI model of a network device serving the terminal device, and the second AI model is an AI model in the cloud.
12. A communication device, wherein, It includes a module or unit for executing the method according to any one of claims 1 to 11.
13. A communication device, wherein, It includes: At least one processor, and the at least one processor is used to execute a computer program or instruction stored in a memory, so that the device executes the method according to any one of claims 1 to 11.
Citation Information
Cited By
Reasoning method and communication apparatus
EP4807631A1
Reasoning method and communication apparatus
WO2025108413A1