Model processing method and apparatus, device, and storage medium
By using the interaction between the terminal device and the network device in a wireless network, a propt template based on input information, type information and/or task information is determined, which solves the problem of inaccurate determination of propt templates in the prior art, and improves the convenience of using large models and the inference effect.
Patent Information
- Application Number
- PCT/CN2024/126562
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-22
- Publication Date
- 2025-05-08
AI Technical Summary
In wireless networks, it is difficult for the prior art to accurately determine the propt template, resulting in low convenience in using large models and the inference effect of model is not guaranteed.
The terminal device sends input information, type information and/or task information to the network device, and requests the network device to determine M first problems based on the first propt template, and determines N first problems through model reasoning and filtering to realize intelligent acquisition of the propt template.
It improves the convenience of model use and model reasoning effect, ensuring the accuracy and applicability of the propt template.
Smart Images

Figure CN2024126562_08052025_PF_FP_ABST
Abstract
Description
Model processing method, device, equipment and storage medium
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on October 30, 2023, with application number 202311433959.X and application name “Model processing method, device, equipment and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of communication technology, and in particular to a method, apparatus, device, and storage medium for model processing. Background Art
[0003] It's widely acknowledged that sixth-generation mobile networks (6G) are moving towards an era of intelligent and inclusive computing. Large models, such as large language models (LLMs), are expected to become the foundational technology for this intelligent and inclusive computing. In the future, large models are expected to be deployed within wireless networks to achieve deep coupling of communication with perception, computing, and control.
[0004] Prompts are a crucial medium for user interaction with large models. Appropriate prompt templates facilitate accurate problem description, thereby improving the reasoning performance of large models. Currently, prompt templates must be downloaded and selected by users from the cloud, making them inconvenient to use with large models. Furthermore, the selection of prompt templates relies heavily on user experience, which doesn't guarantee the applicability of prompt templates and, consequently, the ability to improve the reasoning performance of large models based on the problems described in prompts. Therefore, accurately determining prompt templates in wireless networks and then using them to describe problems, thereby improving model convenience and reasoning performance, is a pressing issue.
[0005] Summary of the Invention
[0006] The embodiments of the present application provide a model processing method, apparatus, device, and storage medium, in order to accurately determine the prompt template in a wireless network, and then describe the problem based on the prompt template to improve the convenience of model use and the model reasoning effect.
[0007] In a first aspect, the present application provides a model processing method. The method may be executed by a terminal device or a component within the terminal device, such as a chip, a chip system, or other functional module capable of invoking and executing a program. For ease of understanding, the following exemplary description uses a terminal device as the execution entity.
[0008] In an example of the first aspect, the method includes: a terminal device sending a first message, the first message including input information, and type information and / or task information, to request a network device to determine N first questions to be input into a first model, and then receiving a second message from the network device, the second message including M first questions determined based on a first prompt template and the input information, wherein N is a positive integer, and the N first questions are part or all of the M first questions, thereby implementing a prompt service in a wireless network. Furthermore, the first prompt template is determined based on the type information and / or task information, thereby implementing intelligent acquisition of prompt templates, improving the convenience of model use, and ensuring the reasoning effect of the model.
[0009] In one possible implementation, the terminal device may send a third message to the first device, where the third message is used to request the first device to perform model reasoning based on N of the M first problems using the first model, thereby improving the convenience of using the model in the wireless network.
[0010] For the purpose of ensuring the accuracy of the prompt template obtained by matching, or ensuring the rationality of the first question, or reducing the complexity of model processing, etc., the first questions can be screened to determine N first questions from M first questions, such as for the purpose of ensuring the accuracy of the prompt template obtained by matching, or ensuring the rationality of the first question, or reducing the complexity of model processing, etc. Exemplarily, the terminal device sends a fourth message, the fourth message is used to request the second device to screen the M first questions using the first model, and receives a fifth message, the fifth message is used to indicate the N first questions among the M first questions.
[0011] In a possible implementation, the N first questions are the first N first questions arranged in descending order of rationality scores after the first model scores the M first questions for rationality, so as to ensure the rationality of the first questions.
[0012] In one possible implementation, when M first questions do not meet the preset conditions, in order to ensure that the first questions are suitable for input into the first model to perform model reasoning to meet the user's model reasoning requirements, further optimization can be performed to obtain questions that meet the preset conditions (such as the second question). Exemplarily, the terminal device sends a sixth message to the network device, the sixth message including first indication information, the first indication information being used to indicate that the M first questions do not meet the preset conditions, and receives a seventh message including a second question, the second question being generated based on a second prompt template and input information, the second prompt template being aggregated based on the M first prompt templates corresponding to the M first questions. The accuracy of the questions generated by the second prompt template obtained by aggregating multiple first prompt templates is higher than the accuracy of the questions generated by using only one first prompt template, that is, a more accurate reasoning result can be obtained after inputting the first model.
[0013] Optionally, the sixth message may further include: input information, and type information and / or task information.
[0014] In one possible implementation, the second question includes a first vector and a second vector. The first vector is obtained based on the input information, and the second vector is obtained based on the vectors mapped by the M first prompt templates. The second question is presented in the form of a vector to improve the accuracy of the question, that is, to improve the reasoning effect after the question is input into the model.
[0015] In one possible implementation, the terminal device sends an eighth message, which is used to request the first device to perform model reasoning based on the second problem through the first model to improve the reasoning effect of the first model on the input information.
[0016] In another example of the first aspect, the method includes: a terminal device sends a first message, the first message is used to request a network device to determine a third question input into the first model, the first message includes: input information and type information, and receives a second message from the network device, the second message includes a third question, the third question is determined based on a vector generated by a prompt generation unit and the input information, the prompt generation unit corresponds to the type information, and provides a solution for realizing intelligent generation of soft prompts in wireless networks.
[0017] In one possible implementation, the third question includes a third vector and a fourth vector. The third vector is obtained based on the input information, and the fourth vector is generated by the prompt generation unit based on the input information. The third question is presented in the form of a vector to improve the accuracy of the question, that is, to improve the reasoning effect after the question is input into the model.
[0018] Based on the two examples of the first aspect above, optionally, the first message is carried in a non-access stratum NAS message or a user plane message.
[0019] Based on the two examples of the first aspect above, optionally, the second message is carried in a NAS message.
[0020] Based on the two examples of the first aspect mentioned above, optionally, the terminal device receives a ninth message, which is used to indicate supported task information and / or supported large models to synchronize network capability information about model processing between the network device and the terminal device.
[0021] In a second aspect, the present application provides a model processing method. The execution subject of this method can be a network device or a component within the network device, such as a chip, a chip system, or other functional module capable of calling and executing a program. Optionally, the network device can be deployed with a functional module that provides a prompt service, or can be replaced with a functional module that provides a prompt service. For ease of understanding, the following exemplary description uses a network device as the execution subject.
[0022] In an example of the second aspect, the method includes: a network device receives a first message, the first message is used to request determination of N first questions input into a first model, N is a positive integer, the first message includes: input information, and type information and / or task information; the network device sends a second message, the second message includes M first questions, the N first questions are part or all of the M first questions, the first question is determined based on a first prompt template and the input information, and the first prompt template is determined based on the type information and / or task information.
[0023] In one possible implementation, the network device matches M first prompt templates in the template library according to the type information and / or task information, and determines the first question corresponding to the first prompt according to each first prompt template in the M first prompt templates and the input information.
[0024] In a possible implementation, the network device receives a fourth message, where the fourth message is used to request that the M first questions be screened using the first model, and sends a fifth message, where the fifth message is used to indicate N first questions among the M first questions.
[0025] In a possible implementation, it also includes: the network device receives a sixth message, the sixth message includes first indication information, the first indication information is used to indicate that the M first questions do not meet the preset conditions, and sends a seventh message, the seventh message includes a second question, the second question is generated based on a second prompt template and input information, and the second prompt template is aggregated based on the M first prompt templates corresponding to the M first questions.
[0026] In a possible implementation, the sixth message further includes: input information, and type information and / or task information.
[0027] In a possible implementation, the method further includes: the network device aggregating M first prompt templates to obtain a second prompt template, and obtaining a second question based on the second prompt template and input information.
[0028] In a possible implementation, the second question includes a first vector and a second vector, the first vector is obtained according to input information, and the second vector is obtained according to vectors mapped by M first prompt templates.
[0029] In another example of the second aspect, the method includes: a network device receives a first message, the first message is used to request determination of a third question input into the first model, the first message includes: input information and type information, and sends a second message, the second message includes a third question, the third question is determined based on a vector generated by a prompt generation unit and the input information, and the prompt generation unit corresponds to the type information.
[0030] In a possible implementation, the method further includes: the network device obtains a prompt generation unit corresponding to the type information, generates a vector according to the input information through the prompt generation unit, and then determines the third question according to the vector and the input information.
[0031] In a possible implementation, the third question includes a third vector and a fourth vector, the third vector is obtained according to the input information, and the fourth vector is generated by the prompt generation unit based on the input information.
[0032] Based on the above two examples of the second aspect, optionally, the first message is carried in a non-access stratum NAS message or a user plane message.
[0033] Based on the above two examples of the second aspect, optionally, the second message is carried in a NAS message.
[0034] In a third aspect, the present application provides a method for model processing. The execution subject of the method may be a policy control function (PCF), such as a PCF module, or a device or network element deployed with a PCF. For ease of understanding, the following exemplary description is based on the PCF as the execution subject.
[0035] The method includes: the PCF sends a ninth message, the ninth message being used to indicate supported task information and / or supported large models, so as to synchronize network capability information related to model processing.
[0036] Optionally, the PCF may obtain network-supported task information and / or supported large models from a network device (such as a functional module providing prompt services in the network device).
[0037] In a fourth aspect, an embodiment of the present application provides a communication device, which includes a processing module and a transceiver module.
[0038] In an example of the fourth aspect, a processing module is used to determine a first message, which is used to request a network device to determine N first questions input into a first model, where N is a positive integer, and the first message includes: input information, and type information and / or task information; a transceiver module is used to send the first message; the transceiver module is also used to receive a second message from the network device, which includes M first questions, the N first questions are part or all of the M first questions, and the first questions are determined based on a first prompt template and the input information, and the first prompt template is determined based on the type information and / or the task information.
[0039] In a possible implementation, the transceiver module is further used to: send a third message, where the third message is used to request the first device to perform model reasoning based on N first questions among the M first questions through the first model.
[0040] In one possible implementation, the transceiver module is also used to: send a fourth message, which is used to request the second device to filter the M first questions through the first model; and receive a fifth message, which is used to indicate N first questions among the M first questions.
[0041] In one possible implementation, the transceiver module is also used to: send a sixth message, the sixth message including first indication information, the first indication information being used to indicate that the M first questions do not meet the preset conditions; receive a seventh message, the seventh message including a second question, the second question being generated based on a second prompt template and the input information, the second prompt template being aggregated based on the M first prompt templates corresponding to the M first questions.
[0042] In a possible implementation, the sixth message further includes: the input information, and the type information and / or the task information.
[0043] In a possible implementation, the second question includes a first vector and a second vector, the first vector is obtained according to the input information, and the second vector is obtained according to the vectors mapped by the M first prompt templates.
[0044] In a possible implementation, the transceiver module is further used to: send an eighth message, where the eighth message is used to request the first device to perform model reasoning based on the second problem through the first model.
[0045] In another example of the fourth aspect, the processing module is used to determine a first message, which is used to request the network device to determine a third question input into the first model, and the first message includes: input information, and type information; the transceiver module is used to send the first message; the transceiver module is also used to receive a second message from the network device, and the second message includes a third question, which is determined based on a vector generated by a prompt generation unit and the input information, and the prompt generation unit corresponds to the type information.
[0046] In a possible implementation, the third question includes a third vector and a fourth vector, the third vector is obtained according to the input information, and the fourth vector is generated by the prompt generation unit based on the input information.
[0047] Based on another example of the fourth aspect above, optionally, the first message is carried in a NAS message or a user plane message.
[0048] Based on another example of the fourth aspect above, optionally, the second message is carried in a NAS message.
[0049] In a possible implementation, the transceiver module is further configured to receive a ninth message, where the ninth message is configured to indicate supported task information and / or supported large models.
[0050] In a fifth aspect, an embodiment of the present application provides a communication device, comprising: a processing module and a transceiver module.
[0051] In an example of the fifth aspect, the transceiver module is used to receive a first message, which is used to request the determination of N first questions input into the first model, where N is a positive integer, and the first message includes: input information, and type information and / or task information; the processing module is used to determine M first questions, where the N first questions are part or all of the M first questions, and the first questions are determined based on a first prompt template and the input information, and the first prompt template is determined based on the type information and / or the task information; the transceiver module is also used to send a second message, which includes M first questions.
[0052] In one possible implementation, the processing module is specifically used to: match M first prompt templates in the template library according to the type information and / or the task information; and determine the first question corresponding to the first prompt according to each first prompt template in the M first prompt templates and the input information.
[0053] In one possible implementation, the transceiver module is further used to: receive a fourth message, which is used to request that the M first questions be screened through the first model; and send a fifth message, which is used to indicate N first questions among the M first questions.
[0054] In one possible implementation, the transceiver module is also used to: receive a sixth message, the sixth message including first indication information, the first indication information being used to indicate that the M first questions do not meet preset conditions; and send a seventh message, the seventh message including a second question, the second question being generated based on a second prompt template and the input information, the second prompt template being aggregated based on the M first prompt templates corresponding to the M first questions.
[0055] In a possible implementation, the sixth message further includes: the input information, and the type information and / or the task information.
[0056] In a possible implementation, the processing module is further configured to: aggregate the M first prompt templates to obtain a second prompt template; and obtain the second question according to the second prompt template and the input information.
[0057] In a possible implementation, the second question includes a first vector and a second vector, the first vector is obtained according to the input information, and the second vector is obtained according to the vectors mapped by the M first prompt templates.
[0058] In another example of the fifth aspect, the transceiver module is used to receive a first message, which is used to request determination of a third question input into the first model, and the first message includes: input information and type information; the processing module is used to determine the third question, which is determined based on a vector generated by a prompt generation unit and the input information, and the prompt generation unit corresponds to the type information; the transceiver module is also used to send a second message, and the second message includes the third question.
[0059] In a possible implementation, the processing module is specifically configured to: obtain a prompt generation unit corresponding to the type information; generate a vector according to the input information through the prompt generation unit; and determine the third question according to the vector and the input information.
[0060] In a possible implementation, the third question includes a third vector and a fourth vector, the third vector is obtained according to the input information, and the fourth vector is generated by the prompt generation unit based on the input information.
[0061] In the two examples of the fifth aspect above, optionally, the first message is carried in a NAS message or a user plane message.
[0062] In the two examples of the fifth aspect above, optionally, the second message is carried in a NAS message.
[0063] In a sixth aspect, the present application provides a model processing device, including: a processing module for determining a ninth message, the ninth message being used to indicate supported task information and / or supported large models; and a transceiver module for sending the ninth message.
[0064] In a seventh aspect, an embodiment of the present application provides a communication device, comprising: a processor, configured to execute the method of the first aspect, the second aspect, the third aspect or each possible implementation method by running a computer program or through a logic circuit.
[0065] In a possible implementation manner, the communication device further includes: a memory, wherein the memory is used to store the computer program.
[0066] In a possible implementation, the communication device further includes: a communication interface, where the communication interface is used to input and / or output signals.
[0067] In an eighth aspect, an embodiment of the present application provides a chip, comprising: a processor for calling and executing computer instructions from a memory, so that a device equipped with the chip executes a method as in the first aspect, the second aspect, the third aspect, or any possible implementation method.
[0068] In a ninth aspect, an embodiment of the present application provides a communication system, comprising: a communication device for executing the method in the first aspect or each possible implementation, and a communication device for executing the method in the second aspect or each possible implementation.
[0069] In a tenth aspect, an embodiment of the present application provides a computer-readable storage medium for storing computer program instructions, which enables a computer to execute a method as in the first aspect, the second aspect, the third aspect, or any possible implementation.
[0070] In an eleventh aspect, an embodiment of the present application provides a computer program, which enables a computer to execute the method in the first aspect, the second aspect, the third aspect or each possible implementation manner.
[0071] In a twelfth aspect, an embodiment of the present application provides a computer program product, comprising computer program instructions, which enable a computer to execute a method as in the first aspect, the second aspect, the third aspect, or any possible implementation manner.
[0072] The beneficial effects of the contents of the above-mentioned second to twelfth aspects and each possible implementation method can be referred to the beneficial effects brought about by the above-mentioned first aspect and each possible implementation method of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] FIG1 shows a possible, non-limiting system schematic.
[0074] FIG2 is a schematic flowchart of a prompt service provided in an embodiment of the present application.
[0075] FIG3 is a schematic diagram of an interactive flow of a model processing method provided in an embodiment of the present application.
[0076] FIG4 is a schematic diagram of an interactive flow of a model processing method provided in an embodiment of the present application.
[0077] FIG5 is a schematic diagram of an interactive flow of a model processing method provided in an embodiment of the present application.
[0078] FIG6 is a schematic diagram of an interactive flow of a model processing method provided in an embodiment of the present application.
[0079] FIG7 is a schematic diagram of an interactive process of model processing provided in an embodiment of the present application.
[0080] FIG8 is a schematic diagram of an interactive process of model processing provided in an embodiment of the present application.
[0081] FIG9 is a schematic diagram of an interactive process of model processing provided in an embodiment of the present application.
[0082] FIG10 is a schematic diagram of an interactive process for synchronizing capability information related to model processing provided in an embodiment of the present application.
[0083] FIG11 is a schematic block diagram of a communication device provided in an embodiment of the present application.
[0084] FIG12 is another schematic block diagram of a communication device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0085] The technical solution in this application will be described below with reference to the accompanying drawings.
[0086] Figure 1 shows a possible, non-limiting system diagram. As shown in Figure 1, the communication system 10 includes a radio access network (RAN) 100 and a core network (CN) 200. The RAN 100 includes at least one RAN node (such as 110a and 110b in Figure 1, collectively referred to as 110) and at least one terminal (such as 120a-120j in Figure 1, collectively referred to as 120). The RAN 100 may also include other RAN nodes, such as wireless relay equipment and / or wireless backhaul equipment (not shown in Figure 1). The terminal 120 is connected to the RAN node 110 wirelessly. The RAN node 110 is connected to the core network 200 wirelessly or by wire. The core network equipment in the core network 200 and the RAN node 110 in the RAN 100 can be different physical devices, or they can be the same physical device that integrates the core network logical functions and the radio access network logical functions.
[0087] The RAN 100 may be a cellular system related to the Third Generation Partnership Project (3GPP), such as a 4G, 5G, or 6G mobile communication system. The RAN 100 may also be an open access network (O-RAN or ORAN), a cloud radio access network (CRAN), or a wireless fidelity (WiFi) system. The RAN 100 may also be a communication system that integrates two or more of the above systems.
[0088] The RAN node 110, which may also sometimes be referred to as access network equipment, RAN entity or access node, etc., constitutes a part of the communication system to help terminals achieve wireless access. The multiple RAN nodes 110 in the communication system 10 may be nodes of the same type or nodes of different types. In some scenarios, the roles of the RAN node 110 and the terminal 120 are relative. For example, the network element 120i in Figure 1 may be a helicopter or a drone, which may be configured as a mobile base station. For the terminal 120j that accesses the RAN 100 through the network element 120i, the network element 120i is a base station; but for the base station 110a, the network element 120i is a terminal. The RAN node 110 and the terminal 120 are sometimes referred to as communication devices. For example, the network elements 110a and 110b in Figure 1 may be understood as communication devices with base station functions, and the network elements 120a-120j may be understood as communication devices with terminal functions.
[0089] In one possible scenario, a RAN node may be a base station, an evolved NodeB (eNodeB), an access point (AP), a transmission reception point (TRP), a next generation NodeB (gNB), a next generation base station in a sixth generation (6G) mobile communication system, a base station in a future mobile communication system, or an access node in a WiFi system. A RAN node may be a macro base station (such as 110a in FIG1 ), a micro base station or an indoor station (such as 110b in FIG1 ), a relay node or a donor node, or a wireless controller in a CRAN scenario. Optionally, a RAN node may also be a server, a wearable device, a vehicle or an onboard device. For example, an access network device in vehicle to everything (V2X) technology may be a road side unit (RSU). All or part of the functions of the RAN node in this application may also be implemented by software functions running on hardware, or by virtualized functions instantiated on a platform (such as a cloud platform). The RAN node in this application may also be a logical node, a logical module or software that can implement all or part of the RAN node functions.
[0090] In another possible scenario, multiple RAN nodes collaborate to assist the terminal in achieving wireless access, and different RAN nodes respectively implement part of the functions of the base station. For example, the RAN node can be a centralized unit (CU), a distributed unit (DU), a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU). The CU and DU can be set separately, or they can be included in the same network element, such as a baseband unit (BBU). The RU can be included in a radio frequency device or radio frequency unit, such as a remote radio unit (RRU), an active antenna unit (AAU), or a remote radio head (RRH).
[0091] A terminal may also be called terminal equipment, user equipment (UE), access terminal, subscriber unit, subscriber station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device.
[0092] The terminal device may be a device that provides voice / data connectivity to users, such as a handheld device or vehicle-mounted device with wireless connection function. At present, some examples of terminals include: mobile phones, tablet computers, computers with wireless transceiver functions (such as laptops, PDAs, etc.), drones, customer-premises equipment (CPE), smart point of sale (POS) machines, mobile internet devices (MID), virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), and so on. assistant, PDA), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, vehicle-mounted devices, wearable devices, terminal devices in a 5G network or terminal devices in a system evolved after 5G, etc.
[0093] To facilitate understanding of the embodiments of the present application, the terms involved are first explained exemplarily.
[0094] A large model, such as a large language model (LLM), can refer to a machine learning model with a large parameter scale (e.g., more than 10 million), which can be applied to fields such as natural language processing, computer vision, and speech recognition. This application only uses the application scenario of a large model as an example, but does not limit the type and scale of the model used in prompt applications.
[0095] Prompt engineering is a core technology for the efficient operation of large models. Prompts can provide clear, concise, and targeted prompts to optimize the model's output. In prompt engineering, user-entered questions can be converted into new questions based on prompt templates to facilitate model understanding and optimize the model's output. The user-entered questions can also be referred to as the description of the problem, the input of the problem, or input information (such as input). Compared to new questions formed based on the model, the user-entered questions can also be referred to as the original question, the original problem description, the original input, the original problem input, etc.
[0096] For example, in a sentiment classification task, the user input question is "I like this movie," and the prompt template is "[X]. All in all, it's a [Z] movie." The "[X]" in the prompt template is the replacement position of the question input by the user, and "[Z]" is the position of the word to be predicted by the model. The new question obtained based on the prompt template conversion is "I like this movie. All in all, it's a [Z] movie." Based on the new question, the network model can predict the word at the [Z] position, such as "interesting" or "touching." Furthermore, based on the new question, the network model can predict the sentiment classification of the word at the [Z] position, such as a 0.9 probability for positive sentiment and a 0.1 probability for negative sentiment.
[0097] As shown in Figure 2, the cloud currently stores a prompt template library. Users download various prompt templates from the library and select an appropriate prompt template based on their experience. The selected prompt template is then used to translate the user's question into a new question. This new question is then fed into a large cloud-based model, which performs model inference on the input question. The user then downloads the inference results from the cloud. In this case, users need to manually download and select prompt templates, making the model less convenient to use and unable to guarantee model inference performance.
[0098] At present, in order to achieve deep coupling of communication, perception, computing, control, etc. of large models within wireless networks, large models can be deployed in wireless networks. Large models can be deployed independently or in a distributed manner, and this application does not limit this. For example, when a large model is deployed independently, it can be deployed in the core network 200 or the access network 100; for another example, when a large model is deployed in a distributed manner, it can be distributed in the core network 200 and the access network 100, or distributed in the core network 200 and the cloud (such as a cloud server). At present, there is a lack of effective technical solutions for how to obtain prompt templates in wireless networks. Therefore, the technical problem to be solved by this application is to realize the intelligent acquisition of prompt templates through information (including signaling and / or data) transmission and processing between devices in a wireless network, so as to improve the convenience of model use and ensure the model reasoning effect.
[0099] Prompts can include hard prompts and soft prompts. Hard prompts are templates based on fixed prompt words, which form new questions by adding fixed prompt words to the input sentence. For example, in the example of the sentiment classification task mentioned above, the prompt template is the template of the hard prompt; soft prompts replace the fixed prompt word template with a learnable vector. In other words, the prompt words in the template of the hard prompt are readable, and the prompt words in the template of the soft prompt are vectorized, such as the soft prompt is a continuous vector that can be recognized by the model. Studies have shown that using an implicit, trained, and model-optimized soft prompt instead of a hard prompt can achieve better results. For example, when the same large model processes the same task, the Bert-base model scores 43.3 using the hard prompt template and 48.3 using the soft prompt template; the Bert-large model scores 39.4 using the hard prompt and 50.6 using the soft prompt. However, in wireless networks, there is currently a lack of effective technical solutions for how to generate soft prompts, which is another technical problem that this application hopes to solve.
[0100] The model processing method provided in the embodiment of the present application will be described below with reference to the accompanying drawings.
[0101] It should be understood that the following description is merely for ease of understanding and explanation, and the method provided in the embodiments of the present application is described using the interaction between a terminal device and a network device as an example. The terminal device may be, for example, the terminal 120 in FIG. 1 , and the network device may be, for example, the RAN node 110 in FIG. 1 , or a device in the core network 200.
[0102] It should also be understood that this should not constitute any limitation on the execution subject of the method provided in this application. As long as it is possible to execute the method provided in the embodiment of this application by running a program having the code of the method provided in the embodiment of this application, it can serve as the execution subject of the method provided in the embodiment of this application. For example, the above-mentioned terminal device can also be replaced by a component in the terminal device, such as a chip, a chip system or other functional module that can call a program and execute the program; the above-mentioned network device can also be replaced by a component in the network device, such as a chip, a chip system or other functional module that can call a program and execute the program.
[0103] In some embodiments, the network device may be deployed with a prompt service functional module to provide auxiliary services for the large model. The prompt service functional module may be, for example, a prompt logic function (PLF) module, which may be a hardware module, a software module, or a combination of hardware and software. The interaction between the terminal device and the network device may be implemented based on the interaction between the terminal device and the prompt service functional module. PLF is merely an exemplary name for the prompt service functional module and is not limited in this application.
[0104] Optionally, the PLF module can be deployed in the core network (such as the core network 200 in Figure 1) or the RAN (such as the RAN100 in Figure 1). When the PLF module is deployed in the RAN, it can be deployed in the CU. This application does not limit this.
[0105] Optionally, this application does not limit the number of PLF modules deployed in a wireless network. When multiple PLF modules are deployed, each PLF module can correspond to at least one model, each PLF module can be used to perform at least one task, and each PLF module can be used to reason about at least one category of problems.
[0106] Figure 3 is a schematic diagram of an interactive flow of a model processing method provided by an embodiment of the present application. In conjunction with Figure 3, the method 300 includes some or all of the following processes.
[0107] At step S310 , a terminal device sends a first message to a network device, requesting the network device to determine N first questions to be input into a first model, where N is a positive integer. The first message includes input information, type information, and / or task information. In response, the network device receives the first message sent by the terminal device.
[0108] S320: The network device sends a second message to the terminal device. The second message includes M first questions, where the N first questions are part or all of the M first questions. The first questions are determined based on a first prompt template and input information, where the first prompt template is determined based on type information and / or task information. Accordingly, the terminal device receives the second message from the network device.
[0109] It should be noted that the N first questions may be inputs to the first model. When the terminal device requests the network device to determine the first question to be input to the first model, the number of first questions may not be indicated. The first question may be a new question converted by the prompt model in the aforementioned example.
[0110] As mentioned above, the input information (input) may include questions input by the user, such as questions input by the user through a terminal device.
[0111] Type information (type) can refer to the type of user questions (e.g., input), such as classification questions, regression questions, etc. For example, in a sentiment classification model, it can include text classification questions (textCLS); text-span classification questions (text-span CLS), where text-span CLS refers to asking about the sentiment of a subject when a sentence contains sentiments under several different themes; text-pair natural language inference questions (text-pair CLS); sequence tagging questions; and text generation questions. It should be understood that the type information included in the first message can include one or more of the above examples.
[0112] Task information (task) can refer to the category of tasks that the model needs to perform, such as sentiment classification tasks, topic classification tasks, code generation tasks, image generation tasks, text generation tasks, etc. Taking the sentiment classification model as an example, task information may include sentiment classification tasks (sentiment), in which the model can infer whether the text is positive (or positive) or negative (or negative); topic classification tasks (topics), in which the model can infer whether the text is a related topic; intention recognition tasks (intention), in which the model can infer the user's intention and prepare a response based on the intention; aspect sentiment tasks (aspect sentiment); natural language inference (NLI) tasks, in which the model infers whether there is a causal relationship between input sentences; named entity recognition (NER) tasks; summarization tasks; text translation tasks, etc. The task information included in the first message may include one or more of the above examples.
[0113] It should be understood that each type of information may correspond to one or more task information, and similarly, each task information may correspond to one or more type information, which is not limited in this application. Each type of information and a task under that type may correspond to one or more prompt templates.
[0114] The type information and task information may be input by the user through the terminal device, or the type information and task information may be identified by the terminal device. For example, the terminal device may determine the type information and / or task information based on the question input by the user.
[0115] To achieve intelligent determination of prompt templates, in an embodiment of the present application, the first prompt template used to determine the first question can be determined based on the type information and / or task information in the first message. Exemplarily, the network device can determine M corresponding first prompt templates based on the type information and / or task information in the first message, and then determine the first question based on each prompt template in the M first prompt templates and the input information, thereby obtaining M first questions corresponding to the M first prompt templates.
[0116] In some embodiments, a plurality of prompt templates may be stored in a wireless network, and the plurality of prompt templates may be stored in the form of a template library. Optionally, the template library may be deployed in a core network. The prompt template library may include a first correspondence, and the first correspondence includes a correspondence between type information, task information, and prompt templates. As shown in Table 1 below, in the template library, when the type information is "text classification problem" and the task information is "emotion classification task", the prompt template is "[X] It is a [Z] movie"; when the type information is "text fragment classification problem" and the task information is "emotional task", the prompt template is "[X] How is the service? [Z]", and so on.
[0117] Table 1
[0118] The first question can be determined by the network device based on the input information and the first prompt template. Referring to Table 1, when the input information is "I like this movie.", the type information is "text classification problem", and the task information is "emotional classification task", the network device matches the M first prompt templates obtained in the template library based on the type information and task information, including "[X] It is a [Z] movie". Then the first question can be "I like this movie. It is a [Z] movie". After the first question is input into the first model for model inference, the result obtained may include "very good, excellent..."
[0119] In some embodiments, the first prompt template may also be determined based on information of the first model, such as an identifier (ID) of the first model, execution parameters of the first model, a type of the first model, and the like. Optionally, the first message may include information of the first model. The information of the model may have a corresponding relationship with the prompt template, and the corresponding relationship may be preset or preconfigured, so that the network device may determine the first prompt template corresponding to the first model in the corresponding relationship based on the information of the first model.
[0120] The first model may be any model deployed in the wireless network, such as the aforementioned large model, and this application does not limit this. One or more models may be deployed in the wireless network, and the first model may be a model that meets the task requirements. The first model may be determined by an instruction of a terminal device, or the first model may be determined based on task information and / or type information.
[0121] Therefore, in this embodiment, the terminal device sends a first message to the network device, the first message including input information, type information, and / or task information, requesting the network device to determine N first questions for input into the first model. The terminal device then receives a second message from the network device, the second message including M first questions determined based on the first prompt template and the input information, thereby implementing a prompt service on the wireless network. Furthermore, the first prompt template is determined based on the type information and / or task information, enabling intelligent acquisition of prompt templates, improving the convenience of model use, and ensuring the model's reasoning effect.
[0122] Figure 4 is a schematic diagram of an interactive flow of a model processing method provided by an embodiment of the present application. Referring to Figure 4 , the method 400 may include some or all of the following processes.
[0123] In the method 400 shown in FIG4 , S410 and S420 are respectively the same as or similar to S310 and S320 in the embodiment shown in FIG3 , and are not described again for the sake of brevity.
[0124] In some embodiments, the number M of first questions included in the second message sent by the network device may be equal to the number N of first questions that need to be input into the first model, or in other words, the M first questions included in the second message sent by the network device are all used as inputs to the first model.
[0125] In other embodiments, the first questions may be screened to determine N first questions from M first questions for various possible dimensions, such as to ensure the accuracy of matching the prompt template, or to ensure the rationality of the first questions, or to reduce the complexity of model processing.
[0126] It should be noted that this application does not limit the terminal device's method of selecting N first questions from M first questions. Optionally, if only for the purpose of reducing the complexity of model processing, the terminal device may randomly select N first questions from M first questions.
[0127] As an example, the screening for the first question may be performed as shown in FIG4 , including:
[0128] S430: The terminal device may send a fourth message to the second device, where the fourth message may be used to request the second device to screen the M first questions using the first model. Correspondingly, the network device receives the fourth message from the terminal device.
[0129] S440: The second device sends a fifth message to the terminal device, where the fifth message is used to indicate N first questions among the M first questions.
[0130] Illustratively, the network device may prioritize the M first questions using the first model in response to the request of the fourth message, and select the top N first questions in descending order of priority.
[0131] The first model prioritizing the M first questions can be implemented by scoring the M first questions. For example, the M first questions are scored based on the rationality of the first questions, and the priorities are sorted in descending order of the scores, that is, the priorities are sorted in descending order of rationality. It should be noted that this application does not limit the basis for the first model to score the first questions. For example, the first model can score the first questions based on the frequency of occurrence of the first questions.
[0132] The first model is deployed in the second device. When the first model is distributed, the second device can be deployed with part of the first model. The second device can be a core network device, an access network device, or a cloud device (such as a cloud server).
[0133] 4 , in S450, the terminal device may send a third message to the first device, where the third message is used to request the first device to perform model reasoning based on N first questions among the M first questions using the first model. The N first questions may be obtained after question screening, or the N first questions may be M first questions (M=N).
[0134] The first model is deployed in the first device. When the first model is deployed in a distributed manner, the first device may be deployed with a portion of the first model. The first device may be a core network device, an access network device, or a cloud device (such as a cloud server). The first device and the second device may be the same device or different devices. When the first device and the second device are different devices, the first model is deployed in a distributed manner on the first device and the second device.
[0135] In some embodiments, the M first questions included in the second message may not all meet the preset conditions. The preset conditions may be different for different types of questions. For example, for a text generation task, the quality of the answer output by the first model to the first question can be judged from the perspective of diversity and / or accuracy. When the quality of the answer meets a threshold, it is determined that the first question meets the preset conditions. Diversity refers to whether the answers given by the large model are the same when multiple different questions are designed; accuracy refers to whether the answers generated by the first model are fluent and whether they contain harmful content.
[0136] The following, in conjunction with S530 and S540 in Figure 5, provides an exemplary explanation of how to obtain new questions for input into the first model when the M first questions do not meet the preset conditions. S510 and S520 in Figure 5 are respectively the same or similar to S310 and S320 in Figure 3, and are not described again for the sake of brevity. This embodiment is implemented based on the example shown in Figure 3, but this application is not limited to this. This embodiment can also be implemented based on the example shown in Figure 4. For example, after executing S440 in Figure 4, S530 and S540 in Figure 5 are executed. When S530 is executed after S440, the first indication information can be used to indicate that N of the M first questions do not meet the preset conditions.
[0137] 5 , in S530 , the terminal device sends a sixth message to the network device, where the sixth message includes first indication information, where the first indication information is used to indicate that the M first questions do not meet a preset condition.
[0138] Optionally, the sixth message may further include: input information, and type information and / or task information. The input information included in the sixth message may be the same as the input information included in the first message, the type information included in the sixth message may be the same as the type information included in the first message, and the task information included in the sixth message may be the same as the task information included in the first message. However, in some scenarios, in order to successfully obtain the first question that meets the preset conditions, at least one of the input information, type information, and input information included in the first message may be adjusted to increase the possibility of successfully obtaining the first question.
[0139] In response to the request of the sixth message, the network device may aggregate the first prompt templates corresponding to the M first questions to obtain a second prompt template, and then obtain the second question based on the second prompt template, so that the second question meets the preset conditions. As previously mentioned, the first prompt template corresponding to the first question refers to the first prompt template used when generating (or converting to obtain) the first question.
[0140] Furthermore, in S540, the network device may send a seventh message to the terminal device, the seventh message including the second question. Optionally, in S550, the terminal device may send an eighth message to the first device based on the received second question, the eighth message being used to request the first device to perform model reasoning based on the second question using the first model.
[0141] Exemplarily, the second question may include a first vector and a second vector, where the first vector is obtained based on input information, and the second vector is calculated based on vectors mapped by M first prompt templates. The second vector may be a second prompt template obtained by the network device based on aggregation of the M first prompt templates, i.e., the aggregated second prompt template may be a softprompt.
[0142] It should be understood that, based on experimental data, aggregating multiple prompt templates (such as the first prompt template) can achieve better results than using only one prompt template (such as the first prompt template). Questions converted based on different prompt templates (such as the first prompt template) are input into the first model respectively, and then reasoning is performed separately to obtain reasoning results. Next, appropriate rules are selected to aggregate multiple prompt templates (such as the first prompt template). For example, global averaging, weighted averaging, or voting, distillation, etc. can be used to obtain the aggregated prompt template (such as the second prompt template). As the number of aggregated prompt templates (such as the first prompt template) increases, the accuracy of the question will also increase. The advantages of using multiple prompt templates (such as the first prompt template) for aggregation are reflected in: a) leveraging the advantages of different prompt templates to achieve complementarity; b) more stable performance in downstream tasks.
[0143] Exemplarily, the network device aggregating the M first prompt templates to obtain the second prompt template may include: the network device embedding each of the M first prompt templates into a low-dimensional vector through a first model.
[0144] For example, let's assume the input message is "I like deep learning" and the three first prompt templates are: "From now on, you are a translator," "I hope you can take on the role of English translation, spelling proofreader, and rhetoric improvement," and "Now I'll let you be a translator. Your goal is to translate any language into English naturally and authentically." The embedding vectors obtained for the three first prompt templates are:
[0145] [-0.73510434 2.41941113 -0.75119304 1.22246059 -0.49535954 0.17048515]
[0146] [0.1686593 0.66778136 0.65839084 1.76373291 0.62441869 1.28184983]
[0147] [-1.22016658 0.26831831 0.40529441 -0.7297467 -0.77142339 0.70344466]
[0148] The network device averages multiple vectors to obtain the second vector, such as:
[0149] [-0.59553721 1.1185036 0.10416407 0.75214893 -0.21412141 0.71859322]
[0150] The network device embeds the input information "I like deep learning" through the first model to obtain the first vector. For example:
[0151] [0.45298121 0.33040098 0.20693714 -1.48242658 -0.14846323 0.00864623]
[0152] After combining the second vector with the first vector, we get the second problem. For example:
[0153] [[-0.59553721 1.1185036 0.10416407 0.75214893 -0.21412141 0.71859322]
[0154] [0.45298121 0.33040098 0.20693714 -1.48242658 -0.14846323 0.00864623]]
[0155] When the first vector and the second vector are combined to form the second problem, the order of the first vector and the second vector is not limited in the present application.
[0156] 6 , step 600 may include S610 and S620 .
[0157] S610: The terminal device sends a first message to the network device, the first message being used to request the network device to determine a third question input into the first model, the first message including input information and type information. Accordingly, the network device receives the first message from the terminal device.
[0158] S620, the terminal device receives a second message from the network device, where the second message includes a third question, and the third question is determined based on a vector generated by a prompt generation unit and the input information, where the prompt generation unit corresponds to the type information.
[0159] Similar to the first question in the aforementioned embodiment, the third question can be an input to the first model.
[0160] The network device may generate a third question based on the prompt generating unit in response to the first message.
[0161] Exemplarily, the network device may convert the input information through a prompt generation unit to obtain a corresponding vector, which may be referred to as a softprompt. The prompt generation unit may be determined by the network device based on the type information. Exemplarily, the network device may determine, based on the type information included in the first message, the prompt generation unit corresponding to the type information in a preset correspondence relationship, where the correspondence relationship includes a correspondence between the prompt generation unit and the type information.
[0162] Exemplarily, the third question may include a third vector and a fourth vector, wherein the third vector is obtained based on the input information, and the fourth vector is the vector generated by the prompt generation unit based on the input information. Optionally, the prompt generation unit may be referred to as a prompt generator or a prompt generation model.
[0163] For example, if the input information is "I like deep learning", the fourth vector (also known as softprompt) generated by the prompt generation unit according to the input information can be:
[0164] The network device maps the input information into a third vector, such as embedding the input information "I like deep learning" through the first model to obtain the third vector. For example, it can be:
[0165] After combining the fourth vector and the third vector, we get the third problem. For example:
[0166] [[0.59553721 1.1185036 0.10416407 0.75214893 -0.21412141 0.71859322]
[0167] [0.45298121 0.43050098 0.20693714 -1.48242658 -0.14846323 0.00864623]]
[0168] When the third vector and the fourth vector are combined to form the third problem, the order of the third vector and the fourth vector is not limited.
[0169] Optionally, in an embodiment of the present application, the first message may further include task information, or the type information in the first message may be replaced with task information.
[0170] Based on this, in this embodiment, a terminal device sends a first message to a network device. The first message includes input information and type information, requesting the network device to determine a third question for the first model input. The terminal device then receives a second message from the network device. The second message includes the third question, which is determined based on the vector generated by a prompt generation unit and the input information. The prompt generation unit corresponds to the type information. This provides a solution for intelligently generating soft prompts in wireless networks.
[0171] Below, taking the network device implemented as a PLF or a PLF module as an example, the deployment of the PLF in a wireless network is exemplarily described with reference to FIG. 7 to FIG. 9 .
[0172] FIG7 is a schematic diagram of an interactive process of model processing provided by an embodiment of the present application. The PLF may be a device or network element deployed with the PLF, the TCF may be a device or network element deployed with the TCF, and the first model may be a device or network element deployed with the first model (such as the first device and / or second device described above). Referring to FIG7 , the model processing method may include some or all of the following steps:
[0173] S701: A terminal device sends a task request to the TCF. This task request may be used to request the establishment of a task session (taskID) with the PLF. The task request may be carried in a task-non-access stratum (T-NAS) message. Optionally, the task request may include at least one of task information, a problem category, a task identifier (e.g., a task ID), and task execution parameters. The task execution parameters may include at least one of input information and information about the first model.
[0174] S702: The TCF matches the PLF based on the task request, such as based on at least one of the input information, task information, problem category, task identifier (such as task ID), and information of the first model in the task request, obtains the corresponding PLF, and notifies the PLF.
[0175] S703: The PLF sends a request response to the terminal device, where the request response is used to confirm the request to establish the task session, thereby establishing the task session between the terminal device and the PLF.
[0176] S704: The terminal device sends a first message to the PLF. Optionally, the first message may be sent in the form of a task data packet.
[0177] S705: The PLF obtains M first prompt templates based on the type information and / or task information in the first message through matching.
[0178] S706: The PLF obtains M first questions based on each of the M first prompt templates and the input information.
[0179] S707: The PLF sends M first questions to the terminal device. The M first questions may be carried by a task data packet.
[0180] S708: The terminal device determines to use the M first questions as questions to be input into the first model, or determines N first questions from the M first questions as questions to be input into the first model.
[0181] S709: The terminal device sends the determined M first questions or N first questions to the first model.
[0182] S710: The first model performs model reasoning based on the input first question to obtain a reasoning result.
[0183] S711: The first model sends the inference result to the terminal device.
[0184] Among them, S704 to S710 have been explained in the above example and will not be repeated for the sake of brevity.
[0185] FIG8 is a schematic diagram of an interactive process of model processing provided by an embodiment of the present application. The access and mobility management function (AMF) may be a device or network element deployed with the AMF. The PLF and the first model may refer to the description of the embodiment of FIG7 . In the embodiment shown in FIG8 , the terminal device may forward information to the PLF proxy through the AMF to implement information transmission between the terminal device and the PLF. The model processing method may include some or all of the following steps:
[0186] S801: A terminal device sends a first message to an AMF. The first message may be carried in a NAS message.
[0187] S802: AMF sends a first message to PLF to implement proxy transfer of the first message.
[0188] S803: The PLF obtains M first prompt templates based on the type information and / or task information in the first message through matching.
[0189] S804: The PLF obtains M first questions based on each of the M first prompt templates and the input information.
[0190] S805: The PLF sends M first questions to the terminal device. The M first questions are carried in NAS information.
[0191] S806: The terminal device determines to use M first questions as questions to be input into the first model, or determines N first questions from the M first questions as questions to be input into the first model.
[0192] S807: The terminal device sends the determined M first questions or N first questions to the first model.
[0193] S808: The first model performs model reasoning based on the input first question to obtain a reasoning result.
[0194] S809: The first model sends the inference result to the terminal device.
[0195] Some steps are similar to those in the previous example and are not described again for the sake of brevity.
[0196] Figure 9 is a schematic diagram of an interactive flow of model processing provided by an embodiment of the present application. In the embodiment shown in Figure 9, the first message is carried in a user plane message. The user plane function (UPF) can be a device or network element deployed with the UPF. The PLF and the first model can refer to the description of the embodiment of Figure 7. In the embodiment shown in Figure 9, the terminal device can forward information to the PLF proxy through the UPF to achieve information transmission between the terminal device and the PLF. The model processing method may include some or all of the following steps:
[0197] S901: The terminal device sends a first message to the UPF. The first message may be carried in a user plane message.
[0198] S902: The UPF sends a first message to the PLF to implement proxy forwarding of the first message.
[0199] S903: The PLF obtains M first prompt templates based on the type information and / or task information in the first message through matching.
[0200] S904: The PLF obtains M first questions based on each first prompt template in the M first prompt templates and the input information.
[0201] S905: The PLF sends M first questions to the terminal device. The M first questions are carried in NAS information.
[0202] S906: The terminal device determines to use the M first questions as questions to be input into the first model, or determines N first questions from the M first questions as questions to be input into the first model.
[0203] S907: The terminal device sends the determined M first questions or N first questions to the first model.
[0204] S908: The first model performs model reasoning based on the input first question to obtain a reasoning result.
[0205] S909: The first model sends the inference result to the terminal device.
[0206] Some steps are similar to those in the previous example and are not described again for the sake of brevity.
[0207] In some embodiments, the network device may synchronize network-related model processing capability information to the terminal device. For example, the network device may synchronize supported task information and / or supported models (including large models) to the terminal device, so that the terminal device can implement any of the above embodiments based on the network-related model processing capability. For example, the task information included in the first message sent by the terminal device to the network device should belong to the task type supported by the network. See Figure 10 below.
[0208] FIG10 is a schematic diagram of an interactive process for synchronizing capability information related to model processing according to an embodiment of the present application. The PCF may be a device or network element deployed with the PCF, and the PLF may refer to the description of the embodiment of FIG7 . In conjunction with FIG10 , the method may include the following steps:
[0209] S1001, PCF can obtain network supported task information and / or supported models (including large models) from PLF. Exemplarily, PCF sends a query message to PLF, where the query message is used to query the network supported task information and / or supported models (including large models), and PLF can send a query response to PCF, where the query response includes supported task information and / or supported models (including large models); alternatively, PCF can read network supported task information and / or supported models (including large models) from PLF. Optionally, the number of supported task information can be one or more, and the supported task information can form a list (such as a task type list); similarly, the number of supported models can be one or more, and the supported models can form a list (such as a model list).
[0210] S1002: The PCF sends a ninth message to the terminal device, indicating supported task information and / or supported models (including large models). Optionally, the ninth message may be carried in a broadcast message, i.e., the PCF may send the ninth message via a broadcast. Accordingly, the terminal device receives the ninth message from the PCF.
[0211] Figure 11 is a schematic block diagram of a communication device provided in an embodiment of the present application. The communication device 1100 may be a terminal or a network device, or a device within a terminal or network device, or a device capable of being used in conjunction with a terminal or network device. As shown in Figure 11 , the device 1100 may include a transceiver module 1110 and a processing module 1120.
[0212] Optionally, the communication device 1100 may correspond to the terminal device in the above method embodiment.
[0213] When executing any method embodiment shown in Figure 3, the processing module 1120 can be used to determine a first message, which is used to request the network device to determine N first questions input into the first model, where N is a positive integer, and the first message includes: input information, and type information and / or task information; the transceiver module 1110 can be used to send the first message; the transceiver module 1110 is also used to receive a second message from the network device, which includes M first questions, the N first questions are part or all of the M first questions, and the first question is determined based on a first prompt template and the input information, and the first prompt template is determined based on the type information and / or the task information.
[0214] When executing any method embodiment shown in Figure 6, the processing module 1120 can be used to determine a first message, which is used to request the network device to determine a third question input into the first model, and the first message includes: input information, and type information; the transceiver module 1110 can be used to send the first message; the transceiver module 1110 is also used to receive a second message from the network device, and the second message includes a third question, which is determined based on the vector generated by the prompt generation unit and the input information, and the prompt generation unit corresponds to the type information.
[0215] It should be understood that the specific process executed by each module has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0216] Optionally, the communication device 1100 may correspond to the network device in the above method embodiment.
[0217] When executing any method embodiment shown in Figure 3, the transceiver module 1110 can be used to receive a first message, which is used to request to determine N first questions input into the first model, where N is a positive integer, and the first message includes: input information, and type information and / or task information; the processing module 1120 can be used to determine M first questions, where the N first questions are part or all of the M first questions, and the first questions are determined based on a first prompt template and the input information, and the first prompt template is determined based on the type information and / or the task information; the transceiver module 1110 is also used to send a second message, which includes M first questions.
[0218] When executing any method embodiment shown in Figure 6, the transceiver module 1110 can be used to receive a first message, which is used to request determination of a third question input into the first model, and the first message includes: input information and type information; the processing module 1120 can be used to determine the third question, which is determined based on the vector generated by the prompt generation unit and the input information, and the prompt generation unit corresponds to the type information; the transceiver module 1110 is also used to send a second message, and the second message includes the third question.
[0219] It should be understood that the specific process executed by each module has been described in detail in the above method embodiment, and for the sake of brevity, it will not be repeated here.
[0220] The transceiver module 1110 in the communication device 1100 may be implemented by a transceiver, for example, corresponding to the transceiver 1220 in the communication device 1200 shown in FIG12 . The processing module 1120 in the communication device 1100 may be implemented by at least one processor, for example, corresponding to the processor 1210 in the communication device 1200 shown in FIG12 .
[0221] When the communication device 1100 is a chip or chip system configured in a communication device (such as a terminal device or a network device), the transceiver module 1110 in the communication device 1100 can be implemented through an input / output interface, circuit, etc., and the processing module 1120 in the communication device 1100 can be implemented through a processor, microprocessor or integrated circuit integrated on the chip or chip system.
[0222] Figure 12 is another schematic block diagram of a communication device provided in an embodiment of the present application. As shown in Figure 12, the communication device 1200 may include: a processor 1210. The processor 1210 may be used to execute the method executed by the first communication device or the second communication device in the above method embodiment.
[0223] In some possible implementations, the communication device 1200 may include a transceiver 1220. The transceiver 1220 may communicate with the processor 1210 via an internal connection path. The processor 1210 may control the transceiver 1220 to transmit and / or receive signals.
[0224] In some possible implementations, the communication device 1200 may include a memory 1230. The memory 1230 may communicate with the processor 1210 via an internal connection path. The memory 1230 and the processor 1210 may be integrated or provided separately. The memory 1230 may also be a memory external to the device. The memory 1230 is used to store instructions, and the processor 1210 is used to execute the instructions stored in the memory 1230 to perform the method in the above method embodiment.
[0225] It should be understood that the communication device 1200 may correspond to the terminal device or network device in the above-mentioned method embodiment, and may be used to execute the various steps and / or processes performed by the first communication device or the second communication device in the above-mentioned method embodiment. Optionally, the memory 1230 may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. The memory 1230 may be a separate device or integrated into the processor 1210. The processor 1210 may be used to execute the instructions stored in the memory 1230, and when the processor 1210 executes the instructions stored in the memory, the processor 1210 is used to execute the various steps and / or processes of the above-mentioned method embodiment corresponding to the terminal device or network device.
[0226] Optionally, the communication device 1200 is the terminal device in the previous embodiment.
[0227] Optionally, the communication device 1200 is the network device in the previous embodiment.
[0228] The transceiver 1220 may include a transmitter and a receiver. The transceiver 1220 may further include an antenna, which may be one or more. The processor 1210, memory 1230, and transceiver 1220 may be integrated on different chips. For example, the processor 1210 and memory 1230 may be integrated in a baseband chip, and the transceiver 1220 may be integrated in a radio frequency chip. The processor 1210, memory 1230, and transceiver 1220 may also be integrated on the same chip. This application does not limit this.
[0229] Optionally, the communication device 1200 is a component configured in a terminal device, such as a chip, a chip system, etc.
[0230] Optionally, the communication device 1200 is a component configured in a network device, such as a chip, a chip system, etc.
[0231] The transceiver 1220 may also be a communication interface, such as an input / output interface, a circuit, etc. The transceiver 1220 , the processor 1210 , and the memory 1230 may all be integrated into the same chip, such as a baseband chip.
[0232] The present application also provides a processing device, including at least one processor, which executes a computer program or logic circuit to cause the processing device to execute the method executed by the first communication device or the second communication device in the above method embodiment. The processing device may also include a memory for storing the computer program.
[0233] The present application also provides a processing device including a processor and an input / output interface. The input / output interface is coupled to the processor. The input / output interface is used to input and / or output information. The information includes at least one of instructions and data. The processor is configured to execute a computer program to cause the processing device to perform the method performed by the terminal device or network device in the above-described method embodiment.
[0234] The present application also provides a processing device including a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the processing device executes the method executed by the terminal device or network device in the above method embodiment.
[0235] It should be understood that the processing device may be one or more chips. For example, the processing device may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD), or other integrated chips.
[0236] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.
[0237] It should be noted that the processor in the embodiments of the present application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0238] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0239] According to the method provided in the embodiment of the present application, the present application also provides a computer program product, which includes: a computer program or a set of instructions, which, when the computer program or a set of instructions is run on a computer, enables the computer to execute the method executed by the terminal device or network device in the above method embodiment.
[0240] According to the method provided in the embodiment of the present application, the present application also provides a computer-readable storage medium, which stores a program. When the program runs on a computer, the computer executes the method executed by the terminal device or network device in the above method embodiment.
[0241] According to the method provided in the embodiment of the present application, the present application also provides a communication system, which may include the aforementioned terminal device or network device.
[0242] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0243] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0244] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for model processing, characterized in that: Applied to terminal equipment, including: Sending a first message, where the first message is used to request the network device to determine N first questions input into a first model, where N is a positive integer, and the first message includes input information, and type information and / or task information; Receive a second message from the network device, the second message including M first questions, the N first questions being part or all of the M first questions, the first questions being determined based on a first prompt template and the input information, and the first prompt template being determined based on the type information and / or the task information.
2. The method according to claim 1, characterized in that Also includes: A third message is sent, where the third message is used to request the first device to perform model reasoning based on the N first questions among the M first questions through the first model.
3. The method according to claim 1 or 2, characterized in that: Also includes: Sending a fourth message, where the fourth message is used to request the second device to screen the M first questions through the first model; A fifth message is received, where the fifth message is used to indicate the N first questions among the M first questions.
4. The method according to any one of claims 1 to 3, characterized in that: Also includes: Sending a sixth message, where the sixth message includes first indication information, where the first indication information is used to indicate that the M first questions do not meet a preset condition; A seventh message is received, where the seventh message includes a second question, where the second question is generated based on a second prompt template and the input information, where the second prompt template is aggregated based on M first prompt templates corresponding to the M first questions.
5. The method according to claim 4, characterized in that The sixth message also includes the input information, and the type information and / or the task information.
6. The method according to claim 4 or 5, characterized in that: The second question includes a first vector and a second vector, wherein the first vector is obtained according to the input information, and the second vector is obtained according to the vectors mapped by the M first prompt templates.
7. The method according to any one of claims 4 to 6, characterized in that: Also includes: An eighth message is sent, where the eighth message is used to request the first device to perform model reasoning based on the second problem through the first model.
8. A method for model processing, characterized in that: Applied to terminal equipment, including: Sending a first message, the first message is used to request the network device to determine a third question input into the first model, the first message comprising: input information, and type information; A second message is received from the network device, the second message comprising a third question, the third question is determined based on a vector generated by a prompt generation unit and the input information, the prompt generation unit corresponding to the type information.
9. The method according to claim 8, characterized in that The third question includes a third vector and a fourth vector, the third vector is obtained according to the input information, and the fourth vector is generated by the prompt generation unit based on the input information.
10. The method according to any one of claims 1 to 9, characterized in that: The first message is carried in a non-access layer NAS message or a user plane message.
11. The method according to any one of claims 1 to 10, characterized in that: The second message is carried in a NAS message.
12. The method according to any one of claims 1 to 11, characterized in that: Also includes: A ninth message is received, where the ninth message is used to indicate supported task information and / or supported large models.
13. A method for model processing, characterized in that: Applied to network equipment, including: Receive a first message, the first message is used to request to determine N first questions input into a first model, N is a positive integer, the first message includes: input information, and type information and / or task information; Send a second message, the second message including M first questions, the N first questions are part or all of the M first questions, the first questions are determined based on a first prompt template and the input information, the first prompt template The determination is based on the type information and / or the task information.
14. The method according to claim 13, characterized in that Also includes: According to the type information and / or the task information, M first prompt templates are obtained by matching in a template library; According to each first prompt template in the M first prompt templates and the input information, a first question corresponding to the first prompt is determined.
15. The method according to claim 13 or 14, characterized in that receiving a fourth message, wherein the fourth message is used to request to screen the M first questions through the first model; A fifth message is sent, where the fifth message is used to indicate N first questions among the M first questions.
16. The method according to any one of claims 13 to 15, characterized in that Also includes: receiving a sixth message, the sixth message including first indication information, the first indication information being used to indicate that the M first questions do not meet a preset condition; A seventh message is sent, where the seventh message includes a second question, where the second question is generated based on a second prompt template and the input information, where the second prompt template is aggregated based on M first prompt templates corresponding to the M first questions.
17. The method according to claim 16, characterized in that The sixth message further includes: the input information, and the type information and / or the task information.
18. The method according to claim 16 or 17, characterized in that Also includes: Aggregating the M first prompt templates to obtain a second prompt template; The second question is obtained according to the second prompt template and the input information.
19. The method according to any one of claims 16 to 18, characterized in that The second question includes a first vector and a second vector, wherein the first vector is obtained according to the input information, and the second vector is obtained according to the vectors mapped by the M first prompt templates.
20. A method for model processing, characterized in that: Applied to network equipment, including: receiving a first message, the first message being used to request determination of a third question input into the first model, the first message comprising: input information and type information; A second message is sent, wherein the second message includes a third question, wherein the third question is determined based on a vector generated by a prompt generation unit and the input information, and the prompt generation unit corresponds to the type information.
21. The method according to claim 20, characterized in that Also includes: Obtain a prompt generation unit corresponding to the type information; Generate a vector according to the input information by the prompt generating unit; The third question is determined according to the vector and the input information.
22. The method according to claim 20 or 21, characterized in that The third question includes a third vector and a fourth vector, the third vector is obtained according to the input information, and the fourth vector is generated by the prompt generation unit based on the input information.
23. The method according to any one of claims 13 to 22, characterized in that The first message is carried in a non-access layer NAS message or a user plane message.
24. The method according to any one of claims 13 to 23, characterized in that The second message is carried in a NAS message.
25. A communication device, characterized in that: The method comprises a module for executing the method as claimed in any one of claims 1 to 12, or comprises a module for executing the method as claimed in any one of claims 13 to 24.
26. A communication device, characterized in that: include: A processor, wherein the processor is configured to execute the method according to any one of claims 1 to 24 by running a computer program or by a logic circuit.
27. The device according to claim 26, characterized in that Also included is a memory for storing the computer program.
28. The device according to claim 26 or 27, characterized in that Also included is a communication interface for inputting and / or outputting signals.
29. A communication system, characterized in that: include: A communication device for executing the method according to any one of claims 1 to 12, and a communication device for executing the method according to any one of claims 13 to 24.
30. A computer-readable storage medium, characterized in that: Used to store computer program instructions, the computer program causing a computer to execute the method according to any one of claims 1 to 24.
31. A computer program product, characterized in that The method comprises computer program instructions which cause a computer to execute the method as claimed in any one of claims 1 to 24.
Citation Information
Patent Citations
Question answering method and device
CN105989084A
Question and answer response method and device, equipment and storage medium
CN111506723A
Message analysis using a machine learning model
US20190199656A1
Multi-system-based intelligent question answering method and apparatus, and device
US20220391426A1
Cited By
Cooperative reasoning resource allocation method and system for large language model in wireless network
CN120378960A