Reasoning method and communication apparatus
The first AI model obtains the second prompt word based on the first prompt word provided by the user and the local information of the first network device, and supplements the context information, solving the problems of low inference accuracy and large delay of the AI model, achieving more efficient user interaction and more accurate inference results.
Patent Information
- Application Number
- PCT/CN2024/133761
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
When users interact with AI models, users may not provide a comprehensive problem description, resulting in low inference accuracy of AI models and requiring multiple interactions to clearly locate user questions, resulting in large delays and poor user experience.
The first prompt word from the terminal device is received through the first AI model, and the second prompt word is obtained based on the first prompt word and the local information of the first network device, supplement the context information describing the first prompt word, reduce the number of interactions with the user, reduce the inference delay, and improve the inference accuracy of the AI model.
It has achieved the reduction of the number of interactions with users, reduced inference delay, improved the inference accuracy of the AI model, improved the user experience, and met the user's needs for service quality.
Smart Images

Figure CN2024133761_30052025_PF_FP_ABST
Abstract
Description
Reasoning method and communication device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 24, 2023, with application number 202311600465.6 and application name “Inference Method and Communication Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of communication technology, and more specifically, to an inference method and a communication device. Background Art
[0003] Artificial intelligence (AI) has been widely used in the field of communications technology. Currently, when users ask questions based on some AI models (such as large language models (LLMs)), they may not provide a comprehensive description of their questions. Consequently, when the question prompt is relatively short, the AI model's inference accuracy is low. To achieve precise service, the AI model must clearly identify the user's question, requiring multiple interactions between the user and the AI model. This results in significant delays before the AI model infers a satisfactory answer, ultimately leading to a poor user experience. Summary of the Invention
[0004] This application provides a reasoning method that enables the AI model to provide users with more accurate responses, thereby improving the user experience.
[0005] In a first aspect, a method of reasoning is provided. The method can be executed by a first AI model, or can be executed by a component of the first AI model (such as a chip or circuit), without limitation.
[0006] The method includes: a first AI model receives a first prompt word from a terminal device, where the first prompt word is used to request a corresponding reply based on the first prompt word; the first AI model obtains a second prompt word based on the first prompt word and local information of the first network device, where the second prompt word is used to describe information corresponding to at least one parameter related to the first prompt word in the local information of the first network device, where the local information of the first network device is characteristic information of a cell covered by the first network device, and the first network device is a network device with a geographical range related to the first prompt word; the first AI model sends a third prompt word to the second AI model, where the second AI model is an AI model that infers a reply corresponding to the first prompt word, and the third prompt word includes the first prompt word and the second prompt word, or the first AI model sends a first reply result to the terminal device based on the third prompt word.
[0007] Based on the above technical solution, the first AI model can supplement the context information corresponding to the first prompt word based on the local information of the first network device. Afterwards, the first AI model can use the prompt word obtained after supplementing the context information for reasoning, or assist the second AI model in reasoning. This can not only reduce the number of interactions with the user and reduce the reasoning delay, but also improve the reasoning accuracy of the AI model, ensure the user's demand for service quality, and enhance the user experience.
[0008] In combination with the first aspect, in certain implementations of the first aspect, the first AI model obtains a fourth prompt word based on the first prompt word, and the fourth prompt word is used to describe at least one parameter related to the first prompt word.
[0009] In the above technical solution, the first AI model uses the received first prompt word as input, which can explore the deeper needs behind the user's prompt word. For example, if the first prompt word is "Please recommend a library", the first AI model can infer from the word "library" in the first prompt word that the user will mainly be indoors and may use wireless networks for operations such as online book browsing and retrieval, library navigation, and electronic cloud note taking. Therefore, it can derive and generate the fourth prompt word "Take wireless service quality into consideration."
[0010] In combination with the first aspect, in certain implementations of the first aspect, the third prompt word further includes a fourth prompt word.
[0011] In combination with the first aspect, in certain implementations of the first aspect, the first prompt word does not include geographic scope information or the geographic scope information contained in the first prompt word does not meet the geographic scope accuracy required for reasoning. Before the first AI model obtains the second prompt word, the method also includes: the first AI model sends a fifth prompt word to the terminal device, the fifth prompt word is used to prompt the input of geographic scope information related to the first prompt word, or the fifth prompt word includes a sixth prompt word and multiple alternative geographic scope information, the sixth prompt word is used to prompt the selection of geographic scope information related to the first prompt word from multiple alternative geographic scope information; the first AI model receives a seventh prompt word from the terminal device, the seventh prompt word is used to indicate the geographic scope information related to the first prompt word, and the first AI model obtains local information of the first network device based on the seventh prompt word.
[0012] Based on the above technical solution, the first AI model can obtain local data of the network device corresponding to the corresponding area based on the seventh prompt word, and supplement the context information of the first prompt word based on the obtained local data of the network device.
[0013] In combination with the first aspect, in certain implementations of the first aspect, the first AI model sends a third prompt word to the second AI model, and the method also includes: the first AI model receives a second reply result from the second AI model, the second reply result being a reply result obtained by the second model based on reasoning based on the third prompt word; the first AI model verifies that the second reply result is a reasonable reply result based on local information of the first network device; and the first AI model sends the second reply result to the terminal device.
[0014] In combination with the first aspect, in certain implementations of the first aspect, the method also includes: the first AI model receives a third reply result from the second AI model, the third reply result being a reply result inferred by the second AI model based on the third prompt word; the first AI model verifies that the third reply result is an unreasonable reply result based on local information of the first network device, and the first AI model sends an eighth prompt word to the second AI model, the eighth prompt word including the third prompt word and the ninth prompt word, and the ninth prompt word is used to prompt the re-inference of the reply result based on the third prompt word.
[0015] In the above technical solution, the first AI model can use the local information obtained from the first network device to verify the inference results of the second AI model, ensuring that the final reply sent to the user is a reasonable reply, and avoiding the second AI model obtaining an inappropriate reply result due to hallucinations, thereby reducing the user experience.
[0016] In combination with the first aspect, in certain implementations of the first aspect, before the first AI model sends the first result to the terminal device based on the third prompt word, the method also includes: the first AI model verifies that the first result is a reasonable response result based on local information of the first network device.
[0017] In combination with the first aspect, in certain implementations of the first aspect, before the first AI model sends a first reply result to the terminal device based on the third prompt word, the method also includes: the first AI model infers a fourth reply result based on the third prompt word; the first AI model verifies that the fourth reply result is an unreasonable reply result based on the local information of the first network device.
[0018] In the above technical solution, the first AI model can use the local information obtained from the first network device to verify its own reasoning results, ensuring that the reply finally sent to the user is a reasonable reply, and avoiding the first AI model obtaining an inappropriate reply result due to hallucinations and reducing the user experience.
[0019] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the first AI model saves a first sample in a local sample library, and the first sample includes a third prompt word and a final response result of the third prompt word.
[0020] In combination with the first aspect, in certain implementations of the first aspect, the first AI model obtains the second prompt word based on the first prompt word and local information of the first network device, including: the first AI model obtains the second prompt word based on the first prompt word, local information of the first network device and a tenth prompt word, and the tenth prompt word is used to describe at least one sample in the sample library that is related to the first prompt word.
[0021] In the above solution, the user's historical examples are added before the first prompt word, and the context information of the user's prompt word is enriched based on the user's historical examples, so that the AI model of the inference response result can also provide users with accurate personalized services based on information related to user habits or preferences during reasoning, thereby further improving the user experience.
[0022] In combination with the first aspect, in certain implementations of the first aspect, the first AI model is an AI model that infers the response corresponding to the first prompt word, and the first AI model is an AI model in the cloud or an AI model of a network device serving a terminal device, or the second AI model is an AI model that infers the response corresponding to the first prompt word, and the first AI model is an AI model of a network device serving a terminal device, and the second AI model is an AI model in the cloud.
[0023] In an exemplary embodiment, the first AI model and the second AI model are large language models (LLMs).
[0024] In a second aspect, a communication device is provided, the device being configured to execute the method provided in the first aspect. Specifically, the device may include units and / or modules, such as a processing unit and / or a communication unit, for executing the method in any aspect or any possible implementation of the first aspect.
[0025] In one implementation, the device is a first AI model. When the device is the first AI model, the communication unit may be a transceiver or an input / output interface; the processing unit may be at least one processor. Optionally, the transceiver may be a transceiver circuit. Optionally, the input / output interface may be an input / output circuit.
[0026] In another implementation, the device is a chip, chip system, or circuit used in the first AI model. When the device is a chip, chip system, or circuit used in the first AI model, the communication unit may be an input / output interface, interface circuit, output circuit, input circuit, pin, or related circuit on the chip, chip system, or circuit; and the processing unit may be at least one processor, processing circuit, or logic circuit.
[0027] In a third aspect, a communication device is provided, comprising: at least one processor, the at least one processor being coupled to at least one memory, the at least one memory being used to store computer programs or instructions, and the at least one processor being used to call and run the computer program or instructions from the at least one memory, so that the communication device executes the method in the first aspect or any possible implementation of the first aspect.
[0028] In one implementation, the device is a first AI model.
[0029] In another implementation, the device is a chip, a chip system, or a circuit used in the first AI model.
[0030] According to a fourth aspect, a processor is provided for executing the method provided in the first aspect.
[0031] For the operations such as sending and acquiring / receiving involved in the processor, unless otherwise specified, or if they do not conflict with their actual functions or internal logic in the relevant descriptions, they can be understood as processor output, reception, input and other operations, and can also be understood as sending and receiving operations performed by the radio frequency circuit and antenna. This application does not limit this.
[0032] In a fifth aspect, a computer-readable storage medium is provided, which stores a program code for execution by a device, wherein the program code includes a method for executing the above-mentioned first aspect or any possible implementation of the first aspect.
[0033] In a sixth aspect, a computer program product comprising instructions is provided, which, when run on a computer, enables the computer to execute the method in the above-mentioned first aspect and any possible implementation of the first aspect.
[0034] In the seventh aspect, a chip is provided, which includes a processor and a communication interface. The processor reads instructions stored in a memory through the communication interface and executes the method in the above-mentioned first aspect or any possible implementation of the first aspect.
[0035] Optionally, as an implementation method, the chip also includes a memory, in which a computer program or instruction is stored, and the processor is used to execute the computer program or instruction stored in the memory. When the computer program or instruction is executed, the processor is used to execute the method in the above-mentioned first aspect or any possible implementation method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] FIG1 is a schematic diagram of an application framework 100 applicable to a communication system according to an embodiment of the present application.
[0037] FIG2 is a schematic diagram of an application framework 200 applicable to a communication system according to an embodiment of the present application.
[0038] FIG3 is a schematic diagram of a communication system 300 applicable to an embodiment of the present application.
[0039] FIG4 is a schematic diagram of a neuron structure.
[0040] FIG5 is a schematic flowchart of a communication method 500 provided in this application.
[0041] FIG6 is a schematic diagram of a cloud-network collaborative reasoning method 600 proposed in this application.
[0042] FIG7 is a schematic structural diagram of a possible AI model of a network device.
[0043] FIG8 is a schematic block diagram of a communication device 800 provided in an embodiment of the present application.
[0044] FIG9 is a schematic block diagram of a communication device 900 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0046] Before introducing the embodiments of the present application, the following points are first explained.
[0047] 1. In the description of the embodiments of the present application, unless otherwise specified, “multiple” means two or more.
[0048] 2. In the various embodiments of the present application, unless otherwise specified or provided for by logic, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0049] 3. The various numerical numbers involved in this application are only used for the convenience of description and are not used to limit the scope of this application. The size of the serial numbers involved in this application does not mean the order of execution. The execution order of each process should be determined by its function and internal logic. For example, the terms "first", "second", "third", "fourth" and other various terminology labels (if any) in the specification and claims and drawings of this application are used to distinguish similar objects and are not used to limit the size, content, order, timing, priority or importance of multiple objects. For example, the first information and the second information do not represent the difference in the amount of information, content, priority or importance.
[0050] 4. The terms "comprise", "include", "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed but may include other steps or units not explicitly listed or inherent to such process, method, product or apparatus.
[0051] 5. In each embodiment of the present application, "network element A sends information A to network element B" can be understood as the destination end of the information A or the intermediate network element in the transmission path between the destination end and the network element B, which may include directly or indirectly sending information to network element B. "Network element B receives information A from network element A" can be understood as the source end of the information A or the intermediate network element in the transmission path between the source end and the network element A, which may include directly or indirectly receiving information from network element A. The information may be processed as necessary between the source end and the destination end of the information transmission, such as format changes, etc., but the destination end can understand the valid information from the source end. Similar expressions in this application can be understood similarly and will not be elaborated here.
[0052] In other words, sending and receiving can be performed between devices, for example, between terminal device #1 and terminal device #2, or can be performed within a device, for example, sending or receiving between components, modules, chips, software modules or hardware modules within the device through a bus, traces or interface.
[0053] 6. In the embodiments of the present application, indications include direct indications (also called explicit indications) and implicit indications. Direct indication of information A refers to including information A; implicit indication of information A refers to indicating information A through the correspondence between information A and information B and the direct indication of information B. The correspondence between information A and information B can be predefined, pre-stored, pre-burned, or pre-configured.
[0054] 7. In the embodiments of the present application, information C is used to determine information D, which includes information D being determined solely based on information C, as well as information D being determined based on information C and other information. Furthermore, information C can also be used to determine information D indirectly, for example, where information D is determined based on information E, and information E is determined based on information C.
[0055] 8. "Storage" or "saving" in the embodiments of this application may refer to storage in one or more memories. The one or more memories may be provided separately or integrated into an encoder or decoder, a processor, or a communication device. The one or more memories may also be provided in part separately and in part integrated into a decoder, a processor, or a communication device. The type of memory may be any form of storage medium and is not limited in this application.
[0056] 9. The “protocol” involved in the embodiments of the present application may refer to a standard protocol in the field of communications, for example, it may include a fourth generation (4G) network / fifth generation (5G) network protocol, a new radio (NR) protocol, and related protocols used in future communication systems. This application does not limit this.
[0057] 10. The dotted arrows or boxes in the schematic diagrams in the accompanying drawings of this application specification represent optional steps or optional modules.
[0058] The technical solutions provided in this application can be applied to various communication systems, such as: fifth generation (5G) or new radio (NR) systems, long term evolution (LTE) systems, LTE frequency division duplex (FDD) systems, LTE time division duplex (TDD) systems, wireless local area networks (WLAN) systems, satellite communication systems, future communication systems, such as sixth generation (6G) mobile communication systems, or a fusion system of multiple systems. The technical solutions provided in this application can also be applied to device to device (D2D) communication, vehicle-to-everything (V2X) communication, machine to machine (M2M) communication, machine type communication (MTC), and Internet of Things (IoT) communication systems or other communication systems.
[0059] A device in a communication system can send signals to or receive signals from another device. These signals may include information, signaling, or data. The term "device" can also be replaced by an entity, network entity, communication device, communication module, node, communication node, and the like. This disclosure uses devices as examples for description. For example, a communication system may include at least one terminal device and at least one network device. A network device can send downlink signals to a terminal device, and / or a terminal device can send uplink signals to a network device.
[0060] In an embodiment of the present application, the terminal device may also be referred to as user equipment (UE), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device.
[0061] The terminal device may be a device that provides voice / data, such as a handheld device or vehicle-mounted device with a wireless connection function. At present, some examples of terminals are: mobile phones, tablet computers, laptop computers, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, wearable devices, terminal devices in 5G networks or future evolved public land mobile communication networks (PLMNs). The terminal equipment in the network (PLMN), etc., is not limited to this in the embodiments of the present application.
[0062] By way of example and not limitation, in an embodiment of the present application, the terminal device may also be a portion of a network device used to implement terminal device functions. For example, the network device may be an integrated access and backhaul (IAB) node. The IAB node integrates a mobile terminal (MT) and a distributed unit (DU), or an MT and a base station (BS), where the BS includes a CU and a DU. When the IAB node faces its parent node, it can be considered a terminal. In this case, the IAB node plays the role of the MT.
[0063] As an example and not a limitation, in the embodiment of the present application, the terminal device may also be a wearable device. Wearable devices may also be called wearable smart devices, which are a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not only hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are fully functional, large in size, and can achieve complete or partial functions without relying on smartphones, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.
[0064] In the embodiments of the present application, the device for realizing the function of the terminal device can be a terminal device, or a device capable of supporting the terminal device to realize the function, such as a chip system, which can be installed in the terminal device or used in combination with the terminal device. In the embodiments of the present application, the chip system can be composed of a chip, or it can include a chip and other discrete devices. In the embodiments of the present application, only the terminal device is used as an example for description, and the embodiments of the present application are not limited to the solutions of the embodiments of the present application.
[0065] The network device in the embodiments of the present application may be a device for communicating with a terminal device, and may also be referred to as an access network device or a radio access network device. For example, the network device may be a base station. The network device in the embodiments of the present application may refer to a radio access network (RAN) node (or device) that connects a terminal device to a wireless network. The base station can broadly cover various names as follows, or be replaced with the following names, such as: NodeB, evolved NodeB (eNB), next generation NodeB (gNB), relay station, IAB node (such as the BS functional part in the IAB node), access point, transmission point (transmitting and receiving point, TRP), transmitting point (transmitting point, TP), master station, auxiliary station, multi-standard radio (motor slide retainer, MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), radio unit (RU), positioning node, etc. The base station can be a macro base station, a micro base station, a relay node, a donor node or the like, or a combination thereof. The base station can also refer to a communication module, a modem or a chip that is provided in the aforementioned device or apparatus. The base station can also be a mobile switching center and a device that performs the base station function in D2D, V2X, and M2M communications, a network side device in a 6G network, a device that performs the base station function in future communication systems, and the like. The base station can support networks with the same or different access technologies. Optionally, the RAN node can also be a server, a wearable device, a vehicle or an on-board device, and the like. For example, the access network device in the vehicle to everything (V2X) technology can be a road side unit (RSU). The embodiments of the present application do not limit the specific technology and specific device form adopted by the network equipment.
[0066] In some deployments, the network devices mentioned in the embodiments of the present application may include a CU, a DU, or both a CU and a DU, or a control plane CU node (central unit-control plane (CU-CP)), a user plane CU node (central unit-user plane (CU-UP)), and a DU node. For example, the network devices may include a gNB-CU-CP, a gNB-CU-UP, and a gNB-DU.
[0067] In some deployments, multiple RAN nodes collaborate to assist terminals in achieving wireless access, with different RAN nodes implementing portions of the base station's functionality. For example, a RAN node can be a CU, DU, CU-CP, CU-UP, or RU. The CU and DU can be separate or included in the same network element, such as the BBU. The RU can be included in a radio frequency device or radio unit, such as an RRU, AAU, or RRH.
[0068] The RAN node may support one or more types of fronthaul interfaces, with different fronthaul interfaces corresponding to DUs and RUs with different functions. If the fronthaul interface between the DU and the RU is a common public radio interface (CPRI), the DU is configured to implement one or more baseband functions, and the RU is configured to implement one or more radio frequency functions. If the fronthaul interface between the DU and the RU is another type of interface, relative to the CPRI, some of the downlink and / or uplink baseband functions, such as precoding, digital beamforming (BF), or one or more of inverse fast Fourier transform (IFFT) / cyclic prefix (CP) for downlink, are moved from the DU to the RU for implementation; and for uplink, one or more of digital beamforming (BF), or fast Fourier transform (FFT) / cyclic prefix (CP) removal, are moved from the DU to the RU for implementation. In one possible implementation, the interface may be an enhanced common public radio interface (eCPRI). In the eCPRI architecture, the division between the DU and RU is different, corresponding to different types (category, Cat) of eCPRI, such as eCPRI Cat A, B, C, D, E, and F.
[0069] Taking eCPRI Cat A as an example, for downlink transmission, based on layer mapping, the DU is configured to implement layer mapping and one or more functions preceding it (i.e., one or more of coding, rate matching, scrambling, modulation, and layer mapping). Other functions after layer mapping (e.g., RE mapping, digital beamforming (BF), or one or more of inverse fast Fourier transform (IFFT) / cyclic prefix (CP) addition) are moved to the RU for implementation. For uplink transmission, based on RE demapping, the DU is configured to implement demapping and one or more functions preceding it (i.e., one or more of decoding, rate matching, descrambling, demodulation, inverse discrete Fourier transform (IDFT), channel equalization, and RE demapping). Other functions after demapping (e.g., one or more of digital BF or fast Fourier transform (FFT) / CP removal) are moved to the RU for implementation. It is understandable that for the functional description of DU and RU corresponding to various types of eCPRI, reference can be made to the eCPRI protocol, which will not be described in detail here.
[0070] In one possible design, the processing unit for implementing baseband functions in the BBU is called a baseband high layer (BBH) unit, and the processing unit for implementing baseband functions in the RRU / AAU / RRH is called a baseband low layer (BBL) unit.
[0071] In different systems, CU (or CU-CP and CU-UP), DU or RU may also have different names, but those skilled in the art can understand their meanings. For example, in the ORAN system, CU may also be called O-CU (Open CU), DU may also be called O-DU, CU-CP may also be called O-CU-CP, CU-UP may also be called O-CU-UP, and RU may also be called O-RU. Any unit of CU (or CU-CP, CU-UP), DU and RU in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.
[0072] In the embodiments of the present application, the device for implementing the functions of the network device can be a network device; it can also be a device that can support the network device to implement the functions, such as a chip system, a hardware circuit, a software module, or a hardware circuit and a software module. The device can be installed in the network device or used in conjunction with the network device. In the embodiments of the present application, only the device for implementing the functions of the network device is used as an example to illustrate, and does not constitute a limitation on the solutions of the embodiments of the present application.
[0073] The network device and / or terminal device can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; it can also be deployed on the water surface; it can also be deployed on aircraft, balloons and satellites in the air. The embodiments of this application do not limit the scenarios in which the network device and the terminal device are located. In addition, the terminal device and the network device can be hardware devices, or they can be software functions running on dedicated hardware, software functions running on general-purpose hardware, such as virtualization functions instantiated on a platform (e.g., a cloud platform), or entities including dedicated or general-purpose hardware devices and software functions. This application does not limit the specific forms of the terminal device and the network device.
[0074] In wireless communication networks, such as mobile communication networks, the services supported by the networks are becoming increasingly diverse, and therefore the demands that need to be met are becoming increasingly diverse. For example, the network needs to be able to support ultra-high speeds, ultra-low latency, and / or ultra-large connections. This feature makes network planning, network configuration, and / or resource scheduling increasingly complex. In addition, as network functionality becomes increasingly powerful, such as supporting higher spectrum, supporting high-order multiple input multiple output (MIMO) technology, supporting beamforming, and / or supporting new technologies such as beam management, network energy saving has become a hot research topic. These new demands, new scenarios, and new features have brought unprecedented challenges to network planning, operation and maintenance, and efficient operation. To meet this challenge, artificial intelligence technology can be introduced into wireless communication networks to achieve network intelligence.
[0075] In order to support AI technology in wireless networks, AI nodes may also be introduced into the network.
[0076] Optionally, the AI node can be deployed in one or more of the following locations in the communication system: access network equipment, terminal equipment, or core network equipment. Alternatively, the AI node can be deployed separately, for example, in a location other than any of the above devices, such as a host or cloud server in an over-the-top (OTT) system. The AI node can communicate with other devices in the communication system, such as one or more of the following: network equipment, terminal equipment, or network elements of the core network.
[0077] It is understood that the embodiments of the present application do not limit the number of AI nodes. For example, when there are multiple AI nodes, the multiple AI nodes can be divided based on function, such as different AI nodes are responsible for different functions.
[0078] It is also understood that AI nodes can be independent devices, or integrated into the same device to implement different functions, or can be network elements in hardware devices, or can be software functions running on dedicated hardware, or can be virtualized functions instantiated on a platform (e.g., a cloud platform). This application does not limit the specific form of AI nodes. Among them, AI nodes can be AI network elements or AI modules.
[0079] Figure 1 is a schematic diagram of an application framework 100 in a communication system applicable to an embodiment of the present application. As shown in Figure 1, network elements in the communication system are connected through interfaces (e.g., NG, Xn) or air interfaces. These network element nodes, such as cloud devices, access network nodes (RAN nodes), and one or more devices in the terminal are provided with one or more AI modules (for clarity, only one is shown in Figure 1). The access network node can be a separate RAN node, or it can include multiple RAN nodes, for example, including CU and DU. CU and / or DU can also be provided with one or more AI modules. Optionally, the CU can also be split into CU-CP and CU-UP. One or more AI models are provided in the CU-CP and / or CU-UP.
[0080] The AI module is used to implement the corresponding AI function. The AI modules deployed in different network elements can be the same or different. The AI module model can implement different functions according to different parameter configurations. The AI module model can be configured based on one or more of the following parameters: structural parameters (such as the number of neural network layers, the neural network width, the connection relationship between layers, the weight of neurons, the activation function of neurons, or at least one of the bias in the activation function), input parameters (such as the type of input parameters and / or the dimension of input parameters), or output parameters (such as the type of output parameters and / or the dimension of output parameters). Among them, the bias in the activation function can also be referred to as the bias of the neural network. An AI module can have one or more models. A model can infer an output, which includes one or more parameters. The learning process, training process, or inference process of different models can be deployed in different nodes or devices, or can be deployed in the same node or device.
[0081] In this communication system, the AI model of the access network node exchanges necessary data with the AI model in the cloud and finally feeds back the data to the user. The data can be carried on a physical channel, for example, it can be carried on a physical downlink control channel (PDCCH), a physical downlink shared channel (PDSCH), a physical uplink shared channel (PUSCH) or a physical uplink control channel (PUCCH), etc. For another example, it can be carried on a physical sidelink control channel (PSCCH) or a physical sidelink shared channel (PSSCH).
[0082] In this communication system, the AI module for the cloud can be located in the cloud device or can be separated from the cloud device. Similarly, for the AI model of the access network device, the AI module can be located in the network device or can be separated from the network device. This application does not impose any restrictions on this.
[0083] Figure 2 is a schematic diagram of an application framework 200 in a communication system applicable to an embodiment of the present application. As shown in Figure 2, the communication system includes a RAN intelligent controller (RIC). For example, the RIC is used to implement AI-related functions. The RIC includes a near-real time RIC (near-real time RIC, near-RT RIC) and a non-real time RIC (non-real time RIC, Non-RT RIC). Among them, the non-real-time RIC mainly processes non-real-time information, such as data that is not sensitive to delay, and the delay of the data can be in the order of seconds. The real-time RIC mainly processes near-real-time information, such as data that is relatively sensitive to delay, and the delay of the data is in the order of tens of milliseconds.
[0084] Near real-time RIC is used for model training and reasoning. For example, it is used to train an AI model and use the AI model for reasoning. Near real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (e.g., CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or reasoning data. Optionally, near real-time RIC can deliver the reasoning results to the RAN node and / or terminal. Optionally, the reasoning results can be exchanged between the CU and DU, and / or between the DU and RU. For example, the near real-time RIC delivers the reasoning results to the DU, and the DU sends it to the RU.
[0085] Non-real-time RIC is also used for model training and reasoning. For example, it is used to train AI models and use the models for reasoning. Non-real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (such as CU, CU-CP, CU-UP, DU and / or RU) and / or terminals. This information can be used as training data or reasoning data, and the reasoning results can be submitted to the RAN node and / or terminal. Optionally, the reasoning results can be exchanged between the CU and DU, and / or between the DU and RU. For example, the non-real-time RIC submits the reasoning results to the DU, and the DU sends it to the RU.
[0086] The near-real-time RIC and non-real-time RIC can also be set up as separate network elements. Optionally, the near-real-time RIC and non-real-time RIC can also be part of other devices. For example, the near-real-time RIC is set up in a RAN node (e.g., a CU or DU), while the non-real-time RIC is set up in an OAM, a cloud server, a core network device, or other network devices.
[0087] Figure 3 is a schematic diagram of a communication system 300 applicable to an embodiment of the present application. As shown in Figure 3, the communication system 300 may include at least one network device, for example, a network device 110; the communication system 100 may also include at least one terminal device, for example, a terminal device 120 and a terminal device 130. The network device 110 and the terminal device (such as the terminal device 120 and the terminal device 130) can communicate via a wireless link. The communication devices in the communication system, for example, the network device 110 and the terminal device 120, can communicate via multi-antenna technology. The communication system 300 also includes an AI network element 140. The AI network element 140 is used to perform AI-related operations, such as building a training data set or training an AI model.
[0088] In one possible implementation, network device 110 may send data related to AI model training to AI network element 140, which may construct a training dataset and train the AI model. For example, data related to AI model training may include data reported by terminal devices. AI network element 140 may send AI model-related operation results to network device 110, which may be forwarded to the terminal device via network device 110. For example, AI model-related operation results may include at least one of the following: a trained AI model, model evaluation results, or test results.
[0089] It should be understood that Figure 4 illustrates only the example of a direct connection between the AI network element 140 and the network device 110. In other scenarios, the AI network element 140 can be connected to both the network device 110 and the terminal device. Alternatively, the AI network element 140 can be connected to the network device 110 via a third-party network element. This embodiment of the present application does not limit the connection relationship between the AI network element and other network elements. For example, the AI network element 140 can also be provided as a module in a network device, for example, in the network device 110 shown in Figure 3.
[0090] It should be noted that Figures 1 and 3 are simplified schematic diagrams for ease of understanding. For example, the communication system may also include other devices, such as wireless relay devices and / or wireless backhaul devices, which are not shown in Figures 1 and 3. In actual applications, the communication system shown in Figures 1 and 3 may include multiple network devices and may also include multiple terminal devices. The embodiments of the present application do not limit the number of network devices and terminal devices included in the communication system.
[0091] It should also be noted that the system architecture and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0092] To facilitate understanding of this application, the following is a brief description of some of the technical concepts involved in this application.
[0093] (1) Machine learning (ML) is an important technical approach to achieving AI. ML can be divided into supervised learning, unsupervised learning, and reinforcement learning.
[0094] Deep neural networks (DNNs) are a specific implementation of machine learning (ML). According to the universal approximation theorem, neural networks can theoretically approximate any continuous function, enabling them to learn arbitrary mappings. Traditional communication systems require extensive expert knowledge to design communication modules. However, deep learning communication systems based on DNNs can automatically discover implicit patterns in large datasets and establish mappings between data, achieving performance superior to traditional modeling methods. The concept of DNNs is inspired by the neuronal structure of the brain. Each neuron performs a weighted summation operation on its input values, and this summation is then applied to a nonlinear function to generate an output. Figure 4 illustrates the neuronal structure.
[0095] As shown in Figure 4, Figure 4 is a schematic diagram of a neuron structure. As shown in Figure 4, assume that the input of the neuron is x=[x0,…,x n ], and the weight corresponding to the input is w=[w0,…,w n ], the bias of the weighted sum is b, and the form of nonlinear function can be diversified. An example is the max{0,x} maximum function. Then the effect of the execution of a neuron can be
[0096] (2) AI model: An AI model is an algorithm or computer program that can implement AI functions. The AI model represents the mapping relationship between the input and output of the model. The type of AI model can be a neural network, a linear regression model, a decision tree model, a support vector machine (SVM), a Bayesian network, a Q learning model, or other ML models. The AI model of the embodiment of the present application can be a two-end model or a single-end model.
[0097] (3) Model reasoning: Use the trained model to solve practical problems.
[0098] (4) Natural Language Processing: Natural language processing is an interdisciplinary subfield of computer science and linguistics. It primarily involves giving computers the ability to support and manipulate language and text. The goal is a computer that can "understand" the content of a document, including the contextual nuances of the language within it. The technology can then accurately extract the information and insights contained in the document and categorize and organize the document itself.
[0099] (5) Prompt word: a natural language text that describes the task that AI should perform.
[0100] (6) Context: In linguistics, context refers to the objects or entities surrounding a focal event. It is a framework that surrounds the event and provides resources for its proper interpretation. The meaning of words in natural language processing is usually inferred from the context in which they appear.
[0101] It is understood that the AI model can be implemented as a hardware circuit, software, or a combination of software and hardware, without limitation. Non-limiting examples of software include: program code, program, subroutine, instruction, instruction set, code, code segment, software module, application, or software application.
[0102] Large language models (LLMs) have recently revolutionized the fields of natural language processing (NLP) and AI. LLMs are AI models based on machine learning algorithms and natural language processing techniques, ushering in a new era in text generation, understanding, and interaction. LLMs are characterized by their large size and capacity. They are trained by pre-processing large amounts of text data collected from the internet using AI algorithms. They typically contain tens to billions of weights and are designed to simulate and understand the characteristics and patterns of human language. LLMs build ultra-large neural network models for applications such as information retrieval, sentiment analysis, text generation, chatbots, and conversational AI. They can be used to answer questions on any online platform, allowing visitors to receive meaningful information without having to contact human experts. In response to user questions, LLM-based intelligent agents are AI-driven programs that use natural language processing and AI to infer answers. Unlike most traditional question-answering applications, LLMs' performance relies on their massive databases, which contain text from diverse sources, including books, blogs, and news articles. This massive database enables them to understand information about a wide range of topics, enabling them to provide better answers. It enables machines to understand user-provided prompts and provide useful assistance to the questioner. As LLM reasoning accuracy improves and its application scope expands, more and more users are interacting with this advanced application to get answers to their questions and help them complete their daily tasks.
[0103] Since LLM has a huge number of parameters that require a huge storage space, and its reasoning process requires the use of powerful hardware to achieve fast calculations, the current LLM reasoning solution is often implemented based on cloud computing, that is, cloud-based LLM reasoning. This solution provides users with an LLM application program interface and allows users (individuals and / or organizations) to access the LLM application and obtain reasoning results from any device with an Internet connection. The LLM application program interface is essentially a connection between two devices, giving them the ability to send information back and forth. The cloud provides LLM to the user terminal, and the user interacts with the cloud through the user terminal to obtain AI intelligent services. The following is a brief description of the interaction process between the user and the cloud, which mainly includes the following steps:
[0104] 1) The user sends a question prompt to the network device through the user terminal. For example, the question prompt sent by the user is: "Please recommend a library";
[0105] 2) The base station sends the user's question prompt to the cloud;
[0106] 3) The AI model runs in the cloud and calculates the result corresponding to the user request;
[0107] 4) The cloud feeds back the results to the base station; the base station sends the results fed back by the cloud to the user.
[0108] In order to reduce the latency of AI model services for users, a cloud-network collaborative reasoning solution has been proposed. Specifically, by offloading the cloud-trained LLM to network devices (such as base stations), part of the inference tasks of the cloud-based AI model are offloaded to the network physical devices. This solution can prevent congestion caused by a large number of users accessing the cloud-based AI model to request inference services. In addition, the solution assumes that network devices have sufficient computing power to perform AI model inference and are closer to users than cloud servers. Therefore, users can interact with the AI model of network devices, reducing transmission latency to a certain extent.
[0109] Whether it is LLM reasoning based on the cloud or LLM reasoning based on cloud-network collaboration, users generally do not provide a comprehensive description of the question. When the question prompt is relatively short, the reasoning accuracy is low. If precise service is to be achieved, the AI model needs to clearly locate the user's question, which requires the user to interact with the AI model multiple times. This leads to a long delay in the AI model inferring that the user has obtained a satisfactory answer, which ultimately leads to a poor user experience.
[0110] In view of this, the present application provides a reasoning method that can effectively solve the above technical problems. The method proposed in the present application is described in detail below with reference to FIG5 .
[0111] Fig. 5 is a schematic flow chart of a communication method 500 provided by the present application. The method 500 includes the following steps.
[0112] S510: The terminal device sends a first prompt word to the first AI model, where the first prompt word is used to request a corresponding response to be returned based on the first prompt word. The first AI model receives the first prompt word from the terminal device.
[0113] For example, if a user using a terminal device wants to search for a library, the first prompt word may be "Please recommend a library."
[0114] For example, if a user using a terminal device wants to search for a park in Beijing, the first prompt word may be "Please recommend a park in Beijing".
[0115] S520. The first AI model obtains a second prompt word based on the first prompt word and the local information of the first network device. The second prompt word is used to describe information corresponding to at least one parameter related to the first prompt word in the local information of the first network device. The local information of the first network device is characteristic information of the cell covered by the first network device. The first network device is a network device in the geographical range related to the first prompt word.
[0116] It can be understood that if the first prompt word does not include geographic scope information (the first prompt word can be "Please recommend a library"), or the geographic scope information contained in the first prompt word does not meet the geographic scope accuracy required for reasoning, the first AI model needs to further interact with the user until the geographic scope accuracy required for reasoning is obtained, wherein the geographic scope accuracy required for reasoning can be understood as the first AI model determining that the accuracy of the acquired geographic scope related to the first prompt word is not less than the geographic scope accuracy required for reasoning. For example, the geographic scope accuracy required for reasoning is a city, and the geographic scope accuracy required for reasoning is a city or a district within a city. For another example, the geographic scope accuracy required for reasoning is a district of a city, and the geographic scope accuracy required for reasoning is a district of a city, or a region within a district of a city.
[0117] Optionally, in the above scenario, the method further includes: the first AI model sends a fifth prompt word to the terminal device, and the fifth prompt word is used to further interact with the user to obtain geographic scope information that meets the geographic scope accuracy. Correspondingly, the terminal device receives the fifth prompt word from the first AI model. Afterwards, the user sends a seventh prompt word to the first AI model through the terminal device, and the seventh prompt word is used to describe the geographic scope information related to the first prompt word. Correspondingly, the first AI model receives the seventh prompt word from the terminal device. Afterwards, the first AI model obtains the local information of the network device (i.e., the first network device) corresponding to the corresponding geographic scope based on the seventh prompt word.
[0118] For example, the fifth prompt word can be used to prompt the user to enter geographic range information related to the first prompt word. For example, if the first prompt word is "Please recommend a library," the fifth prompt word can be "Enter the approximate range of the library you wish to recommend, such as Shanghai Pudong New Area, Beijing Haidian District, etc."
[0119] For example, the fifth prompt may also include a sixth prompt and multiple alternative geographic range information. The sixth prompt prompts the user to select a geographic range information related to the first prompt from the multiple alternative geographic range information. For example, if the first prompt is "Please recommend a library," the fifth prompt may be "Please select the approximate range of the library you wish to recommend from the following options." The sixth prompt may include multiple options corresponding to geographic ranges, and the user may select the option corresponding to the geographic range they desire related to the first prompt.
[0120] Then, when the first AI model obtains geographic range information that meets the required inference accuracy, it can obtain a second prompt word based on the local information of the network device corresponding to the obtained geographic range information (i.e., the first network device). The second prompt word is used to describe information corresponding to at least one parameter related to the first prompt word in the local information of the first network device. The following are two implementation methods for obtaining the second prompt word based on the first prompt word and the local information of the first network device.
[0121] In the first implementation, the second prompt can be viewed as the first AI model mining information based on the first prompt (mining at least one parameter related to the first prompt), supplementing this mined information with the local information of the first network device (obtaining information related to the mined at least one parameter from the local data of the first network device), thereby improving the first prompt to better derive the response. The following example illustrates this.
[0122] For example, the first prompt word can be "Please recommend a library", and the parameter related to the first prompt word can be wireless quality data. Assuming that the first AI model interacts with the user and determines that the user needs to recommend a library in Pudong New Area, where the current user is located, the first AI model obtains information related to wireless quality data stored in the local information of the network device (i.e., the first network device) within the Pudong New Area. The information related to wireless quality data may include but is not limited to the following parameters: wireless traffic data, number of active users, and user CQI. For example, the second prompt word can be "The current base station is located in Pudong New Area. The wireless traffic in this area on Monday was a1, and the number of connected users was b1. The wireless traffic on Tuesday was a2, and the number of connected users was b2. ... The future wireless traffic will be x, and the number of connected users will be u."
[0123] In a second implementation, the first AI model infers multiple response results based on the first prompt word. The first AI model then mines information based on the first prompt word (mining at least one parameter related to the first prompt word). The model then supplements the mined information for the multiple response results based on the local information of the first network device (obtaining information related to at least one parameter related to each response result from the local data of the first network device), thereby improving the multiple response results to determine the optimal response result. An example is provided below.
[0124] For example, the first prompt word may be "Please recommend a library". Assume that after the first AI model interacts with the user, it is clear that the user needs to recommend a library in Pudong New Area where the current user is located. The first AI infers multiple Pudong New Area libraries (i.e., multiple reply results). For example, multiple Pudong New Area libraries include "Library A, Library B, and Library C". Afterwards, the first AI model obtains information related to the wireless quality data related to the multiple reply results stored in the local information of the network device (i.e., the first network device) within the Pudong New Area. For example, the second prompt word may be "The current base station is located in Pudong New Area. The wireless traffic in the area where Library A is located on Monday was a11, and the number of connected users was b11. The wireless traffic on Tuesday was a12, and the number of connected users was b12. ..., the future wireless traffic will be x1, and the number of connected users will be u1. The wireless traffic in the area where Library B is located on Monday was a21, and the number of connected users was b21. The wireless traffic on Tuesday was a22, and the number of connected users was b22. ..., the future wireless traffic will be x2, and the number of connected users will be u2. The wireless traffic in the area where Library C is located on Monday was a31, and the number of connected users was b31. The wireless traffic on Tuesday was a32, and the number of connected users was b32. ..., the future wireless traffic will be x3, and the number of connected users will be u3.
[0125] Optionally, before S520, the method further includes: the first AI model obtains a fourth prompt word based on the first prompt word, and the fourth prompt word is used to describe at least one parameter related to the first prompt word. For example, the first prompt word is "Please recommend a library", and the fourth prompt word may be "Take the quality of wireless service into consideration". For example, the first prompt word may be "Please recommend a park", and the fourth prompt word may be "Take the environmental conditions into consideration". For example, the information related to the environmental perception data in the local information of the first network device may include but is not limited to information such as temperature, humidity, air pressure, and wind speed.
[0126] Afterwards, the first AI model obtains the second prompt word based on the first prompt word, the local information of the first network device, and the fourth prompt word.
[0127] It should be understood that this application does not impose any restrictions on whether the first AI model generates the fourth prompt word. The first AI model can also skip generating the fourth prompt word and directly obtain the second prompt word. This application does not impose any restrictions on this.
[0128] It can be understood that in the first implementation method, a supplementary description of the first prompt word is provided based on the local information of the first network device to improve the context information of the first prompt word, so that the AI model that performs subsequent reasoning can better reason about the first prompt word based on the prompt word with the supplemented context information, and obtain a more accurate response result. In the second implementation method, based on the first prompt word, the first AI model first roughly derives multiple response results for the first prompt word. Then, based on the local information of the first network device, a supplementary description of the multiple response results is provided to improve the context information of the corresponding responses, so that the AI model that performs subsequent reasoning can select the best response result from multiple results based on the multiple response results with the supplemented context information.
[0129] For example, the AI model of the subsequent reasoning response result may be the first AI model or another AI model.
[0130] It should be understood that if the AI model of the subsequent reasoning response result is the first AI model, S530 is executed. For example, the first AI model can be an AI model of a network device serving the terminal device, or the second AI model can be an AI model in the cloud.
[0131] It should be understood that if the AI model of the subsequent reasoning response result is the second AI model, S540 is executed. For example, the first AI model can be the AI model of the network device in the cloud-network collaborative reasoning scenario, and the second AI model is the AI model of the cloud in the cloud-network collaborative reasoning scenario.
[0132] For example, the first AI model and the second AI model are LLMs.
[0133] At step S530, the first AI model sends a first reply result to the terminal device, where the first reply result is determined based on a third prompt word, where the third prompt word includes the first prompt word and the second prompt word. In response, the first AI model receives the first reply result from the terminal device.
[0134] For example, the second prompt word may further include a fourth prompt word.
[0135] Based on the above-mentioned implementation method one, for example, if the first prompt word is "Please recommend a library", the third prompt word can be "Please recommend a library, taking wireless quality data into consideration. The current base station is located in Pudong New Area. The wireless traffic in this area on Monday is a1, and the number of connected users is b1. The wireless traffic on Tuesday is a2, and the number of connected users is b2,..., and the future wireless traffic will be x, and the number of connected users will be u." Afterwards, the first AI model can perform inference based on the third prompt word after supplementing the context information, and send the first reply result obtained by inference to the terminal device (user).
[0136] Based on the above-mentioned implementation method 2, for example, if the first prompt word is "Please recommend a library", the third prompt word can be "Please recommend a library, taking the wireless quality data into consideration. The current base station is located in Pudong New Area. The wireless traffic in the area where Library A is located on Monday is a11, and the number of access users is b11. The wireless traffic on Tuesday is a12, and the number of access users is b12,..., the future wireless traffic is x1, and the number of access users is u1. The wireless traffic in the area where Library B is located on Monday is a21, and the number of access users is b21. The wireless traffic on Tuesday is a22, and the number of access users is b22,..., the future wireless traffic is x2, and the number of access users is u2. The wireless traffic in the area where Library C is located on Monday is a31, and the number of access users is b31. The wireless traffic on Tuesday is a32, and the number of access users is b32,..., the future wireless traffic is x3, and the number of access users is u3". Afterwards, the first AI model can perform inference based on multiple reply results after supplementing the context information, and select the best reply result (first reply result) from the multiple reply results and send it to the terminal device (user).
[0137] Optionally, since multiple responses corresponding to the first prompt word have been given in the second implementation, the third prompt word may not include the first prompt word, and this application does not impose any restrictions on this.
[0138] In one possible implementation, the first AI model can use the local information obtained from the first network device to verify its own reasoning results to ensure that the final reply sent to the user is a reasonable reply, thereby avoiding the first AI model obtaining an inappropriate reply due to hallucinations and reducing the user experience. For example, the user wants to recommend a library in Shanghai's Pudong New Area, but the first AI model finally gives a reply of a library in Shanghai's Jing'an District.
[0139] Optionally, before S530, the method further includes: the first AI model verifies the first reply result as a suitable reply result based on local information of the first network device.
[0140] Optionally, before S530, the method further includes: the first AI model inferring a fourth response result based on the third prompt word; and the first AI model verifying that the fourth response result is inappropriate based on local information of the first network device. Thereafter, the first AI model may re-infer a response result based on the third prompt word until the first AI model verifies that the response result is appropriate.
[0141] S540: The first AI model sends a third prompt word to the second AI model. Correspondingly, the second AI model receives the third prompt word from the first AI model.
[0142] For the description of the third prompt word, please refer to the description in S530 and will not be repeated here.
[0143] Optionally, after the second AI model obtains the third prompt word, the method further includes:
[0144] At step S550 , the second AI model sends a response result #1 to the first AI model. The response result #1 is the response result inferred by the second AI model based on the third prompt word. Correspondingly, the first AI model receives the response result #1 from the second AI model.
[0145] S560, the first AI model sends a response result #1 to the terminal device.
[0146] Optionally, in this scenario, the first AI model can also use the local information obtained from the first network device to verify the inference results of the second AI model, ensuring that the final response sent to the user is a reasonable response and preventing the second AI model from receiving inappropriate responses due to hallucinations, which could degrade the user experience. In this implementation, the second AI model is allowed to interact with the first AI model. Compared to independent verification by the second AI model, which is unable to obtain local, unique information from the first network device, this collaborative verification of the inferenced response results can better ensure the validity of the results fed back to the user.
[0147] Then, based on the collaborative verification method proposed above, after the first AI model receives the reply result #1, the method also includes: the first AI model verifies that the reply result #1 is a suitable reply result based on the local information of the first network device, and the first AI model sends the reply result #1 to the terminal device. Correspondingly, the terminal device receives the reply result #1 from the first AI model; or, if the first AI model verifies that the reply result #1 is an inappropriate reply result based on the local information of the first network device, the first AI model does not execute S560, and the first AI model sends the eighth prompt word to the second AI model. The eighth prompt word includes the third prompt word and the ninth prompt word, and the ninth prompt word is used to prompt the re-inference of the reply result based on the third prompt word. Correspondingly, the second AI model receives the eighth prompt word from the first AI model. Afterwards, the second AI model re-infers a reply result #2 based on the third prompt word and sends it to the first AI model. The first AI model continues to verify whether the reply result #2 is a suitable reply result based on the local information of the first network device.
[0148] Based on the above technical solution, the first AI model can use the localized information of the first network device associated with the first prompt word to perform precise inference, or assist the second AI model in performing precise inference, thereby reducing the number of questions posed to the user. Furthermore, by verifying the inference results based on the local data of the first network device, the user's demand for service quality can be further guaranteed on the basis of precise inference, thereby improving the user experience.
[0149] The following continues to describe method 500. In method 500, the first AI model may store a first example corresponding to the final response result of the first prompt word sent to the user terminal in a local example library, where the first example consists of <"first prompt word", "third prompt word", and "final response result of the first prompt word">.
[0150] It can be understood that the local sample library of the first AI model also stores samples corresponding to multiple prompt words previously input by the user terminal. Then, for method 500, in S510, when the first AI model receives the first prompt word from the user terminal, it can select multiple samples related to the first prompt word from the user's historical samples in the sample library and add them before the first prompt word. Then, in S520, the first AI model obtains a second prompt word based on the first prompt word, the local information of the first network device, and the tenth prompt word, wherein the tenth prompt word is used to describe at least one historical sample related to the first prompt word in the user's historical samples contained in the sample library of the first AI model. In addition, the third prompt word used for the first AI model to derive the reply result in S530 and the third prompt word used for the second AI model to derive the reply result in S540 include the tenth prompt word, the first prompt word, and the second prompt word. The subsequent steps are the same and will not be repeated here. This solution enriches the contextual information of the user prompt word based on the user's historical samples, so that the AI model that infers the reply result can also provide users with accurate personalized services based on information related to user habits or preferences during reasoning, thereby further improving the user experience.
[0151] To facilitate understanding of the solution shown in method 500, the following describes the method 500 by taking the cloud-network collaborative reasoning scenario as an example based on the second implementation method in S520.
[0152] Figure 6 is a schematic diagram of a method 600 for cloud-network collaborative reasoning proposed in this application. It should be understood that method 600 is merely an implementation flow for a possible specific scenario of cloud-network collaborative reasoning. The AI model of the network device can be considered the first AI model in method 500, and the AI model in the cloud can be considered the second AI model in method 500. The method includes the following steps.
[0153] In step S601, the user terminal sends prompt word #1 (ie, an example of a first prompt word) to the AI model of the network device. Prompt word #1 is “Please recommend a library.” Correspondingly, the AI model of the network device receives prompt word #1 from the user terminal.
[0154] S602: The AI model of the network device determines that prompt word #1 does not contain geographic range information, and sends prompt word #2 (i.e., an example of the fifth prompt word) to the user terminal. Prompt word #2 is used to further interact with the user to obtain geographic range information that meets the geographic range accuracy.
[0155] 603. The user terminal sends prompt word #3 (i.e., the seventh prompt word) to the AI model of the network device. Prompt word #3 is used to describe the geographical range information related to prompt word #1. Correspondingly, the AI model of the network device receives prompt word #3 from the terminal device.
[0156] For ease of description, the geographical range information indicated by prompt word #3 is taken as an example here for explanation, and area #1 meets the geographical range accuracy required for the inference response result.
[0157] S604: The AI model of the network device infers prompt word #4 (ie, an example of the fourth prompt word) based on prompt word #1.
[0158] For example, if the first AI model locally stores the user's historical examples, the first AI model can select at least one example related to prompt word #1 from the user's historical examples and add it before prompt word #1. If the historical examples are added, prompt word #1 can be considered to include the selected historical examples and prompt word #1.
[0159] The network device's AI model uses prompt word #1 as input to identify the user's underlying needs. For example, based on the word "library," the network device's AI model infers that the user will primarily be indoors and likely use the wireless network for online book browsing and retrieval, library navigation, and electronic cloud note-taking. Therefore, it generates prompt word #4, "Consider wireless service quality."
[0160] Optionally, the AI model of the network device can be a small-capacity long short-term memory (LSTM) model, or a large-capacity model such as a generative pre-trained transformer (GPT) or a bidirectional encoder representation from transformers (BERT). As long as the storage space and computing power of the network device allow, this application does not limit the specific size and scale of the AI model.
[0161] Figure 7 is a schematic structural diagram of a possible AI model of a network device. The model is a GPT model. Optionally, the AI model parameters of the network device are generated by the network device through a certain strategy, for example, they can be generated through pre-training, downloaded from the cloud, or obtained from other third-party entities. As shown in Figure 7, the AI model includes 12 transformation modules (Transformer Block), each of which includes a masked self-attention module, a layer normalization module, a forward neural network module, and a layer normalization module. The user prompt word is learned as the input of the first Transformer Block, and then the output of the first Transformer Block is used as the input of the second Transformer Block to continue learning until the last Transformer Block infers an output.
[0162] S605, the AI model of the network device infers prompt word #5 (i.e., an example of the second prompt word) based on prompt word #4 and local data #1 of network device #1 (i.e., an example of the first network device), where network device #1 is the network device corresponding to area #1.
[0163] For example, prompt word #4 is "Take wireless service quality into consideration", and by calling the specific data related to the wireless service data in data #1, the inferred prompt word #5 can be "The current base station is located in area #1, the wireless traffic in this area on Monday was a1, the number of access users was b1, the wireless traffic on Tuesday was a2, the number of access users was b2,..., the future wireless traffic will be x, and the number of access users will be u".
[0164] Optionally, before the AI model of the network device calls data #1 of network device #1, the AI model of the network device also needs to request the user to authorize it to use the local data of the corresponding network device during reasoning. After the user completes the authorization, the network device can call the local database of the corresponding network device (for example, network device #1) when using the AI model of the network device to perform reasoning.
[0165] S606: The AI model of the network device sends prompt word #6 (i.e., an example of the third prompt word) to the AI model in the cloud. Correspondingly, the AI model in the cloud receives prompt word #6 from the AI model of the network device.
[0166] For example, prompt #6 includes prompt #1, prompt #4, and prompt #5. That is, prompt #6 reads, "Please recommend a library. Taking wireless service quality into consideration, the current base station is located in area #1. On Monday, the wireless traffic in this area was a1, and the number of connected users was b1. On Tuesday, the wireless traffic was a2, and the number of connected users was b2. ..., and the future wireless traffic will be x, and the number of connected users will be u." This shows that the cloud-based AI model provides a supplementary description of prompt #1. The contextual information of prompt #1 is then added to prompt #6, which is ultimately sent to the cloud-based AI model. This approach can help the cloud-based AI model infer more accurate responses.
[0167] S607, the AI model in the cloud infers the response result R1 based on the prompt word #6.
[0168] S608: The AI model in the cloud sends a response result R1 to the AI model in the network device. Correspondingly, the AI model in the network device receives the response result R1 from the AI model in the cloud.
[0169] S609: The AI model of the network device verifies the response result R1 based on data #1 and determines that the response result R1 is an inappropriate response result. Execution then continues with S610.
[0170] At step S610, the AI model on the network device sends prompt #7 (an example of the eighth prompt) to the AI model on the cloud. Prompt #7 includes prompt #6 and prompt #8 (an example of the ninth prompt). Prompt #8 prompts the user to re-infer the response based on prompt #6. For example, prompt #8 is "Please provide a response other than R1." In response, the AI model on the cloud receives prompt #7 from the AI model on the network device.
[0171] S611, the cloud-based AI model infers the response result R2 based on the prompt word #6.
[0172] S612: The AI model in the cloud sends a response result R2 to the AI model in the network device. Correspondingly, the AI model in the network device receives the response result R2 from the AI model in the cloud.
[0173] S613, the AI model of the network device verifies that the response result R2 is a suitable response result based on data #1.
[0174] S614: The AI model of the network device sends a response result R2 to the user terminal. Correspondingly, the user terminal receives the response result R2 from the AI model of the network device.
[0175] It can be understood that method 600 is described in detail using prompt word #1 as an example of "recommended a library". For example, if prompt word #1 is "please recommend a park", then based on this prompt word, the only difference from method 600 is that the contextual information that needs to be supplemented in the derivation process is different. Specifically, the AI model of the network device in S604 infers that the user will mainly be outdoors based on the "park" in prompt word #1, so it derives prompt word #4 "taking environmental conditions into account". The AI model of the network device in S605 calls the specific data related to the environmental perception data in data #1 based on prompt word #4 "taking environmental conditions into account". The inferred prompt word #5 can be "The current base station is located in area #1, and the temperature at 8 o'clock in this area is a1, the humidity is b1,... the temperature at 9 o'clock is a2, and the humidity is b 2, …, in the future, the temperature at 8 o'clock in this area will be x1, and the humidity will be x1, …, the temperature at 9 o'clock will be x2, and the humidity will be x2, …”. In S606, the AI model of the network device sends prompt word #6 to the AI model in the cloud. Prompt word #6 is “Please recommend a park. Taking environmental conditions into consideration, the current base station is located in area #1, and the temperature at 8 o'clock in this area is a1, and the humidity is b1, …, the temperature at 9 o'clock is a2, and the humidity is b2, …, in the future, the temperature at 8 o'clock in this area will be x1, and the humidity will be y1, …, the temperature at 9 o'clock will be x2, and the humidity will be y2, …”. The remaining steps are similar to method 600 and are not repeated here.
[0176] It can be understood that some optional features in the various embodiments of the present application may not depend on other features in certain scenarios, and may also be combined with other features in certain scenarios, without limitation.
[0177] It can also be understood that the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.
[0178] It is also understood that in some of the above embodiments, devices in existing network architectures are mainly used as examples for illustrative purposes, and it should be understood that the embodiments of the present application do not limit the specific form of the devices. For example, devices that can achieve the same functions in the future are applicable to the embodiments of the present application.
[0179] It can also be understood that in the above-mentioned various method embodiments, the methods and operations implemented by the first AI model can also be implemented by components of an AI model (such as a chip or circuit).
[0180] The method provided in the embodiment of the present application is described in detail above with reference to Figures 1 to 7. The above method is mainly introduced from the perspective of the interaction between the first AI model, the first AI model, and the second AI model. It is understandable that in order to implement the above functions, the first AI model includes the corresponding hardware structure and / or software modules for performing each function.
[0181] Those skilled in the art should be aware that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is performed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0182] Below, the communication device provided in the embodiment of the present application is described in detail with reference to Figures 8 and 9. It should be understood that the description of the device embodiment corresponds to the description of the method embodiment. Therefore, for the content that is not described in detail, please refer to the above method embodiment. For the sake of brevity, some content will not be repeated. In the embodiment of the present application, the first AI model can be divided into functional modules according to the above method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation. The following is an example of dividing each functional module corresponding to each function.
[0183] The communication method provided by this application is described in detail above. The communication device provided by this application is described below. In one possible implementation, the device is used to implement the steps or processes corresponding to the first AI model in the above method embodiment.
[0184] Figure 8 is a schematic block diagram of a communication device 800 provided in an embodiment of the present application. As shown in Figure 8, the device 800 may include a communication unit 810 and a processing unit 820. The communication unit 810 can communicate with the outside world, and the processing unit 820 is used to process data. The communication unit 810 may also be referred to as a communication interface or a transceiver unit.
[0185] In one possible design, the device 800 can implement steps or processes corresponding to the execution of the first AI model in the above method embodiment, wherein the processing unit 820 is used to perform processing-related operations of the first AI model in the above method embodiment, and the communication unit 810 is used to perform sending-related operations of the first AI model in the above method embodiment.
[0186] It should be understood that the device 800 here is embodied in the form of a functional unit. The term "unit" here may refer to an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor or a group processor, etc.) and a memory for executing one or more software or firmware programs, a merged logic circuit and / or other suitable components that support the described functions. In an optional example, those skilled in the art will understand that the device 800 can be specifically the first AI model in the above embodiment, and can be used to execute the various processes and / or steps corresponding to the first AI model in the above method embodiment. To avoid repetition, they are not described here.
[0187] The device 800 of each of the above schemes has the function of implementing the corresponding steps performed by the first AI model in the above method. The function can be implemented by hardware, or it can be implemented by hardware executing the corresponding software. The hardware or software includes one or more modules corresponding to the above functions; for example, the communication unit can be replaced by a transceiver (for example, the sending unit in the communication unit can be replaced by a transmitter, and the receiving unit in the communication unit can be replaced by a receiver), and other units, such as the processing unit, can be replaced by a processor to respectively perform the sending and receiving operations and related processing operations in each method embodiment.
[0188] In addition, the above-mentioned communication unit can also be a transceiver circuit (for example, it can include a receiving circuit and a transmitting circuit), and the processing unit can be a processing circuit. In an embodiment of the present application, the device in Figure 8 can be the first AI model in the aforementioned embodiment, or it can be a chip or a chip system, such as a system on chip (SoC). Among them, the communication unit can be an input and output circuit, a communication interface; the processing unit is a processor or microprocessor or integrated circuit integrated on the chip. This is not limited here.
[0189] Figure 9 is a schematic block diagram of a communication device 900 provided in an embodiment of the present application. The device 900 includes a processing circuit 910 and a transceiver circuit 920. The processing circuit 910 and the transceiver circuit 920 communicate with each other via an internal connection path. The processing circuit 910 is used to execute instructions to control the transceiver circuit 920 to send and / or receive signals. The transceiver circuit 920 can also be referred to as a transceiver interface.
[0190] Optionally, the apparatus 900 may further include a memory 930, which communicates with the processing circuit 910 and the transceiver circuit 920 via an internal connection path. The memory 930 is used to store instructions, and the processing circuit 910 can execute the instructions stored in the memory 930. In one possible implementation, the apparatus 900 is used to implement the various processes and steps corresponding to the first AI model in the above method embodiment.
[0191] It should be understood that the device 900 can be specifically the first AI model in the above-mentioned embodiment, or it can be a chip or a chip system. Correspondingly, the transceiver circuit 920 can be the transceiver interface of the chip, which is not limited here. Specifically, the device 900 can be used to execute the various steps and / or processes corresponding to the first AI model in the above-mentioned method embodiment. Optionally, the memory 930 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type. The processing circuit 910 can be used to execute instructions stored in the memory, and when the processing circuit 910 executes the instructions stored in the memory, the processing circuit 910 is used to execute the various steps and / or processes of the above-mentioned method embodiment corresponding to the first AI model.
[0192] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.
[0193] It should be noted that the processing circuit in the embodiments of the present application can be a processing circuit in an integrated circuit chip with signal processing capabilities. During implementation, each step of the above-mentioned method embodiment can be completed by hardware integrated logic circuits in the processing circuit or by software instructions. The above-mentioned processing circuit can be or be included in the following devices: a general-purpose processor, a digital signal processing (DSP), an ASIC, a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The processor in the embodiments of the present application can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processing circuit reads the information in the memory and, in conjunction with its hardware, completes the steps of the above-mentioned method.
[0194] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0195] It should be noted that when the processing circuit is a processor or is included in a processor, and the processor is a general-purpose processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) can be integrated into the processor.
[0196] In addition, the present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are run on a computer, the operations and / or processes performed by the first AI model in each method embodiment of the present application are executed.
[0197] The present application also provides a computer program product, which includes computer program code or instructions. When the computer program code or instructions are run on a computer, the operations and / or processes performed by the first AI model in each method embodiment of the present application are executed.
[0198] In addition, the present application further provides a chip, the chip including a processing circuit. A memory for storing a computer program is provided independently of the chip or is provided within the chip, and the processing circuit is configured to execute the computer program stored in the memory, so that the operations and / or processing performed by the first AI model in any method embodiment are performed.
[0199] Furthermore, the chip may further include a communication interface. The communication interface may be an input / output interface, or an interface circuit, etc. Furthermore, the chip may further include a memory.
[0200] In addition, the present application also provides a communication system, including the first AI model and terminal device in the embodiment of the present application. Optionally, the communication system also includes a second AI model.
[0201] It should also be noted that the memory described herein is intended to comprise, but not be limited to, these and any other suitable types of memory.
[0202] Those skilled in the art will appreciate that the various exemplary units and algorithmic steps described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented using hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application. Those skilled in the art will clearly understand that, for ease of description and brevity, the specific operating processes of the systems, devices, and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units described is merely a logical functional division. In actual implementation, other divisions may be used, such as multiple units or components being combined or integrated into another system, or some features being omitted or not implemented. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interface, or indirect coupling or communication connection between devices or units, which may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0203] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.
[0204] It should be understood that references to "embodiments" throughout this specification mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present application. Therefore, various embodiments throughout this specification do not necessarily refer to the same embodiment. Furthermore, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0205] It should also be understood that in this application, "when", "if" and "if" all mean that the network element will make corresponding processing under certain objective circumstances, which is not a time limit, and does not require the network element to make judgment actions when implementing it, nor does it mean that there are other limitations.
[0206] It should also be understood that, in this application, "at least one" means one or more, and "plurality" means two or more. "At least one item" or similar expressions refers to one or more items, that is, any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c means: a, b, c, a and b, a and c, b and c, or a, b, and c.
[0207] It should also be understood that the term "and / or" in this document simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A alone, including A and B, and B alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects are in an "or" relationship. For example, "A / B" means: A or B.
[0208] It should also be understood that in each embodiment of the present application, "A corresponds to B" means that B is associated with A, and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A, and B can also be determined based on A and / or other information.
Claims
1. A method of reasoning, characterized in that include: The first artificial intelligence AI model receives a first prompt word from a terminal device, where the first prompt word is used to request a corresponding reply to be returned based on the first prompt word; The first AI model acquires a second prompt word based on the first prompt word and local information of the first network device, where the second prompt word is used to describe information corresponding to at least one parameter related to the first prompt word in the local information of the first network device, the local information of the first network device is characteristic information of a cell covered by the first network device, and the first network device is a network device in a geographical range related to the first prompt word; The first AI model sends a third prompt word to the second AI model, the second AI model is an AI model that infers a corresponding answer to the first prompt word, and the third prompt word includes the first prompt word and the second prompt word. or, The first AI model sends a first reply result to the terminal device based on the third prompt word.
2. The method according to claim 1, characterized in that The first AI model obtains a fourth prompt word based on the first prompt word, where the fourth prompt word is used to describe at least one parameter related to the first prompt word.
3. The method according to claim 2, characterized in that The third prompt word also includes the fourth prompt word.
4. The method according to any one of claims 1 to 3, characterized in that The first prompt word does not include geographic scope information or the geographic scope information included in the first prompt word does not meet the geographic scope accuracy required for reasoning. Before the first AI model acquires the second prompt word, the method further includes: The first AI model sends a fifth prompt word to the terminal device, the fifth prompt word is used to prompt input of geographic range information related to the first prompt word, or the fifth prompt word includes a sixth prompt word and multiple candidate geographic range information, the sixth prompt word is used to prompt selection of geographic range information related to the first prompt word from the multiple candidate geographic range information; The first AI model receives a seventh prompt word from the terminal device, where the seventh prompt word is used to indicate geographic range information related to the first prompt word; The first AI model obtains local information of the first network device based on the seventh prompt word.
5. The method according to any one of claims 1 to 4, characterized in that The first AI model sends the third prompt word to the second AI model, and the method further includes: The first AI model receives a second reply result from the second AI model, where the second reply result is a reply result inferred by the second AI model based on the third prompt word; The first AI model verifies, based on local information of the first network device, that the second reply result is a reasonable reply result; The first AI model sends the second reply result to the terminal device.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: The first AI model receives a third reply result from the second AI model, where the third reply result is a reply result inferred by the second AI model based on the third prompt word; The first AI model verifies that the third reply result is an unreasonable reply result based on the local information of the first network device, and the first AI model sends an eighth prompt word to the second AI model, and the eighth prompt word includes the third prompt word and the ninth prompt word, and the ninth prompt word is used to prompt to re-infer the reply result based on the third prompt word.
7. The method according to any one of claims 1 to 6, characterized in that Before the first AI model sends the first result to the terminal device based on the third prompt word, the method further includes: The first AI model verifies that the first result is a reasonable response result based on local information of the first network device.
8. The method according to claim 7, characterized in that Before the first AI model sends a first reply result to the terminal device based on the third prompt word, the method further includes: The first AI model infers a fourth answer result based on the third prompt word; The first AI model verifies that the fourth reply result is an unreasonable reply result based on the local information of the first network device.
9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: The first AI model saves a first sample in a local sample library, where the first sample includes the third prompt word and a final response result of the third prompt word.
10. The method according to claim 9, characterized in that The first AI model acquires a second prompt word based on the first prompt word and local information of the first network device, including: The first AI model obtains the second prompt word based on the first prompt word, local information of the first network device and a tenth prompt word, where the tenth prompt word is used to describe at least one sample in the sample library that is related to the first prompt word.
11. The method according to any one of claims 1 to 10, characterized in that The first AI model is an AI model for inferring a response corresponding to the first prompt word, and the first AI model is an AI model in the cloud or an AI model of a network device serving the terminal device. or, The second AI model is an AI model for inferring the answer corresponding to the first prompt word, the first AI model is an AI model of the network device serving the terminal device, and the second AI model is an AI model in the cloud.
12. A communication device, characterized in that: The method comprises a module or a unit for executing the method according to any one of claims 1 to 11.
13. A communication device, characterized in that: include: At least one processor, wherein the at least one processor is configured to execute a computer program or instruction stored in a memory so that the apparatus performs the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Inference method and communication device
CN120046722A
Video dialogue question and answer data generation method and device, electronic equipment and medium
CN116842216A
Conversation prompt text generation method and device, electronic equipment and storage medium
CN116955547A
Chained natural language interaction method and system driven by large language model
CN116992006A
Method and system for automatically generating a response to a user query
US20180268456A1
Cited By
Cooperative reasoning resource allocation method and system for large language model in wireless network
CN120378960A
Input system, device and method for artificial intelligence generation content of intelligent cabin and storage medium
CN121075321A