A natural language processing method, apparatus and electronic device

CN117114007BActive Publication Date: 2026-09-18TSINGHUA UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311070516.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-24
Publication Date
2026-09-18
Estimated Expiration
2043-08-24

AI Technical Summary

Technical Problem

但是当用户输入的文本不完整时,移动终端就不能准确理解用户的意图,降低了用户的使用体验

Benefits of technology

[0014] In the solutions provided by the first to fourth aspects of this application, upon receiving natural language text input by the user, contextual information about the user's current situation within a preset time period is obtained. The natural language text and contextual information are then processed to obtain device response information for the user's input. Compared to related technologies where incomplete user input prevents accurate understanding of the user's intent, this approach combines the user's input and contextual information to understand the user's expressed intent. Even if the user's input is incomplete or unclear, the user's expressed intent can be understood more accurately. Responding to the user with an accurate understanding of their intent makes the mobile terminal's response more aligned with the user's intent, thus improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117114007B_ABST
    Figure CN117114007B_ABST
Patent Text Reader

Abstract

The application provides a natural language processing method, device and electronic equipment, and the method comprises the following steps: when a natural language text input by a user is received, a mobile terminal acquires context information of a context in which the user is located within a preset time period, wherein the context information comprises time information, location information, an application currently used by the user and motion state information of the user; the natural language text and the context information are processed to obtain device response information for the natural language text input by the user; and the device response information is displayed to the user. According to the natural language processing method, device and electronic equipment provided by the application, the natural language text input by the user and the context information of the context in which the user is located can be combined together, the intention expressed by the user can be understood, and the intention expressed by the user can be more accurately understood even if the natural language input by the user is incomplete and unclear.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and more specifically, to a natural language processing method, apparatus, and electronic device. Background Technology

[0002] Currently, natural language semantic understanding typically involves mobile terminals performing semantic analysis on user-input text, obtaining the semantic understanding result, and then responding to the user accordingly. However, when the user-input text is incomplete, the mobile terminal cannot accurately understand the user's intent, thus degrading the user experience. Summary of the Invention

[0003] To address the aforementioned problems, the purpose of this application is to provide a natural language processing method, apparatus, and electronic device.

[0004] In a first aspect, embodiments of this application provide a natural language processing method, including:

[0005] When the mobile terminal receives natural language text input by the user, it obtains contextual information about the user's current situation within a preset time period. The contextual information includes: time information, location information, the application currently being used by the user, and the user's motion status information.

[0006] The natural language text and the contextual information are processed to obtain device response information for the natural language text input by the user;

[0007] The device's response information is displayed to the user.

[0008] Secondly, embodiments of this application also provide a natural language processing apparatus, including:

[0009] The acquisition module is used to acquire contextual information of the user's current situation within a preset time period when it receives natural language text input by the user. The contextual information includes: time information, location information, the application currently being used by the user, and the user's motion status information.

[0010] The processing module is used to process the natural language text and the contextual information to obtain device response information for the natural language text input by the user;

[0011] The display module is used to show the device's response information to the user.

[0012] Thirdly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the method described in the first aspect above.

[0013] Fourthly, embodiments of this application also provide an electronic device, the electronic device including a memory, a processor and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor using the steps of the method described in the first aspect above.

[0014] In the solutions provided by the first to fourth aspects of this application, upon receiving natural language text input by the user, contextual information about the user's current situation within a preset time period is obtained. The natural language text and contextual information are then processed to obtain device response information for the user's input. Compared to related technologies where incomplete user input prevents accurate understanding of the user's intent, this approach combines the user's input and contextual information to understand the user's expressed intent. Even if the user's input is incomplete or unclear, the user's expressed intent can be understood more accurately. Responding to the user with an accurate understanding of their intent makes the mobile terminal's response more aligned with the user's intent, thus improving the user experience.

[0015] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart of a natural language processing method provided in Embodiment 1 of this application is shown;

[0018] Figure 2 This illustration shows a structural schematic diagram of a natural language processing device provided in Embodiment 2 of this application;

[0019] Figure 3 A schematic diagram of the structure of an electronic device provided in Embodiment 3 of this application is shown. Detailed Implementation

[0020] In the description of this application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.

[0021] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0022] In this application, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0023] Currently, natural language semantic understanding typically involves mobile terminals performing semantic analysis on user-input text, obtaining the semantic understanding result, and then responding to the user accordingly. However, when the user-input text is incomplete, the mobile terminal cannot accurately understand the user's intent, thus degrading the user experience.

[0024] Based on this, the following embodiments of this application propose a natural language processing method, apparatus, and electronic device. Upon receiving natural language text input by a user, the device obtains contextual information about the user's current situation within a preset time period, and processes the natural language text and contextual information to obtain device response information for the user's input. This allows the device to combine the user's input natural language text with the contextual information of the user's current situation to understand the user's expressed intent. Even if the user's input natural language is incomplete or unclear, the device can more accurately understand the user's expressed intent and respond to the user with accurate understanding of the user's intent. This makes the mobile terminal's response to the user more consistent with the user's intent, improving the user experience.

[0025] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0026] Example 1

[0027] See Figure 1 The flowchart shown illustrates a natural language processing method. This embodiment proposes a natural language processing method, including the following specific steps:

[0028] Step 100: When the mobile terminal receives natural language text input by the user, it obtains contextual information about the user's current situation within a preset time period. The contextual information includes: time information, location information, the application currently being used by the user, and the user's motion status information.

[0029] In step 100 above, the mobile terminal uses the sensor installed in the mobile terminal to obtain context information of the user's location within a preset time period. In addition to the time information, location information, the application currently used by the user, and the user's motion status information mentioned above, the context information also includes, but is not limited to: wireless network status information, volume information, and new information received by the application.

[0030] The context information also includes: the sensor identifier of the sensor that acquired the context information.

[0031] The sensor refers to the device and application software in the mobile terminal that can obtain contextual information about the user's current situation.

[0032] The time information is obtained through the mobile terminal's system clock.

[0033] Therefore, the system clock of a mobile terminal is a sensor.

[0034] Location information, which indicates the user's current location, is obtained through the WiFi module, GPS positioning module, or Bluetooth module set in the mobile terminal.

[0035] In a mobile terminal, the WiFi module, GPS positioning module, or Bluetooth module that can obtain location information are sensors.

[0036] The application the user is currently using is obtained using accessibility technology provided by the mobile terminal.

[0037] Therefore, the accessibility technology provided by mobile terminals is the sensor.

[0038] User's motion status information, including but not limited to: stationary, walking, running, and cycling.

[0039] In mobile terminals, the IMU motion sensor, which can acquire information about the user's motion status, is called a sensor.

[0040] Mobile terminals include, but are not limited to: mobile phones, laptops, any system / platform / device with any of the above sensors, and devices capable of acquiring contextual information through mobile terminals.

[0041] Step 102: Process the natural language text and the contextual information to obtain device response information for the natural language text input by the user.

[0042] In step 102 above, in order to obtain device response information for the natural language text input by the user, the following steps (1) to (5) can be performed:

[0043] (1) Obtain multiple pre-set API code texts, each of the multiple API code texts including: API identifier, definition, parameters used and functions implemented;

[0044] (2) Process the API code text and the natural language text using a large language model to determine the API code text involved in the natural language text;

[0045] (3) Process the natural language text and the context information to obtain a list of conditional expressions for key phrases in the context information;

[0046] (4) Merge the API code text involved in the determined natural language text and the list of conditional expressions for key phrases in the context information to obtain a markup language file;

[0047] (5) The obtained markup language file is used as the device response information of the natural language text input by the user, thereby obtaining the device response information of the natural language text input by the user.

[0048] In step (1) above, multiple API code texts are cached in the mobile terminal.

[0049] In one implementation, the API identifier refers to a unique name for the function implemented. Examples include "openAPP" and "setVolume".

[0050] API definition: refers to the number of parameters and their order required to implement a function. For example, "setting volume" requires "audio type" and "target volume value" as parameters in that order.

[0051] API parameters refer to the various parameter requirements needed to implement the function. For example, in "Set Volume," "Audio Type" includes "Media," "Alarm," "Incoming Call," and "Message." The "Target Volume Value" ranges from 0 to 15, where 0 is mute and 15 is the maximum target volume value.

[0052] The API implements a function: a piece of natural language text describing the function it performs. For example, "set volume" means "set the volume of a specific audio type on the device to the target value"; "open application" means "open a target application".

[0053] In step (2) above, the Large Language Model (LLM) is a pre-trained semantic analysis model that runs in the mobile terminal as a user assistant. The specific semantic analysis functions that the Large Language Model can achieve are existing technologies and will not be elaborated here.

[0054] The process of using a large language model to process the API code text and the natural language text to determine the specific API code text involved in the natural language text is existing technology and will not be described in detail here.

[0055] In step (3) above, in order to obtain a list of conditional expressions for key phrases in the contextual information, the following steps (31) to (37) can be performed:

[0056] (31) The natural language text and the contextual information are processed using a large language model to obtain a contextual description text;

[0057] (32) The natural language text, the context information and the context description text are processed using a large language model to obtain the first keyword in the context description text;

[0058] (33) Determine a first list of sensors for the sensors that acquire the context information;

[0059] (34) Obtain a second perceptron list with all perceptrons, and use a large language model to process the first perceptron list, the second perceptron list and the first keyword to obtain the perceptron corresponding to the first keyword.

[0060] (35) Perform generalization processing on the perceptron corresponding to the first keyword to obtain a list of first conditional expressions for the first keyword;

[0061] (36) Process the natural language text input by the user to obtain a list of second conditional expressions for the second keyword in the natural language text;

[0062] (37) Merge the first conditional expression list and the second conditional expression list to obtain the conditional expression list of the key phrases in the context information.

[0063] In step (31) above, semantic analysis of natural language text and contextual information is performed using a large language model to obtain contextual description text. The specific process is existing technology and will not be described in detail here.

[0064] In one implementation, when the preset duration is the time period from March 27, 2023, at aa:bb to March 27, 2023, at cc:dd, the context description text may include, but is not limited to: the time is March 27, 2023, at aa:bb; the user's location is the second teaching building; the user's current movement state is stationary; the currently used application is video playback application B; ... the time is March 27, 2023, at cc:dd; the user's location is the playground; the user's current movement state is walking; the currently used application is social application W; and the information sent and received by the user during the time period from March 27, 2023, at aa:bb to March 27, 2023, at cc:dd.

[0065] The information sent and received as described above includes, but is not limited to, text information and voice information.

[0066] In step (32) above, a large language model is used to perform semantic analysis on the natural language text, the contextual information, and the contextual description text to obtain the first keyword in the contextual description text. The specific process is existing technology and will not be described in detail here.

[0067] In one implementation, the first keyword in the context description text may be, but is not limited to, "meeting" and "exercise".

[0068] In step (33) above, the sensor identifier of the sensor that determines the context information is obtained from the context information, and a first sensor list is formed, thereby determining the first sensor list of the sensor that obtains the context information.

[0069] In step (34) above, a second sensor list containing all the sensors is pre-cached in the mobile terminal.

[0070] The specific process of using a large language model to process the first perceptron list, the second perceptron list, and the first keyword to obtain the perceptron corresponding to the first keyword is existing technology and will not be elaborated here.

[0071] In one implementation, the perceptron corresponding to the first keyword can be represented as follows:

[0072] "Meeting" corresponds to "Meeting-related applications"; "Exercise and Fitness" corresponds to "Detecting the user's exercise status".

[0073] In step (35) above, the specific process of using a large language model to generalize the perceptron corresponding to the first keyword to obtain the list of first conditional expressions for the first keyword is existing technology and will not be described in detail here.

[0074] After generalizing the perceptron corresponding to the first keyword, the resulting list of first conditional expressions for the first keyword can be: "meeting" corresponds to "meeting-related applications" or "location = meeting room"; "exercise and fitness" corresponds to "the user's possible exercise state: running, skipping rope, or cycling".

[0075] In one implementation, the conditional expressions in the conditional expression list can be Boolean expressions.

[0076] As can be seen from the above steps (31) to (35), by generalizing, all contextual information corresponding to the first keyword can be found as much as possible. Even if the natural language text input by the user is vague or incomplete, the contextual information corresponding to the first keyword obtained after generalization can be used to more accurately understand the customer's intent.

[0077] Specifically, in order to obtain a list of second conditional expressions for the second keyword in the natural language text, the above step (36) can be performed by following steps (361) to (363):

[0078] (361) The natural language text and the contextual information are processed using a large language model to obtain the second keyword in the contextual description text;

[0079] (362) Use a large language model to process the first perceptron list, the second perceptron list and the second keyword to obtain the perceptron corresponding to the second keyword;

[0080] (363) Perform generalization processing on the perceptron corresponding to the second keyword to obtain a list of second conditional expressions for the second keyword.

[0081] The specific process of implementing steps (361) to (363) above is similar to the process described in steps (32) to (35) above, and will not be repeated here.

[0082] In step (37) above, merging the first conditional expression list and the second conditional expression list involves using a large language model to perform semantic analysis on the first and second conditional expression lists, removing semantically similar conditional expressions from the first and second conditional expression lists, and merging the conditional expressions in the first and second conditional expression lists after removing semantically similar conditional expressions to obtain a list of conditional expressions for key phrases in the contextual information.

[0083] The specific process of obtaining the markup language file in step (4) above is existing technology and will not be described in detail here.

[0084] Markup language files, including but not limited to: JSON format files, YAML format files, and CSV format files.

[0085] In step (5) above, if the natural language text input by the user is a question raised by the user, then the device response information is the answer to the question raised by the user.

[0086] If the natural language text input by the user requires the mobile terminal to create an execution rule, then the device's response information will be the execution rule created based on the natural language text input by the user.

[0087] After generating the execution rule, the mobile terminal assigns a rule identifier to the newly generated execution rule.

[0088] When the device responds with an execution rule created by the mobile terminal based on the natural language text input by the user, step 102 above further includes the following step (6):

[0089] (6) The markup language file is processed using a large language model to obtain the natural language interpretation text of the execution rule.

[0090] In step (6) above, the specific process of using a large language model to perform semantic analysis on the markup language file to obtain the natural language interpretation text of the execution rule is existing technology and will not be described in detail here.

[0091] In one implementation, the natural language interpretation text of the enforcement rule is a natural language interpretation of the machine-readable formatted code. For example, if the enforcement rule is a rule to mute while running on the track, then the natural language interpretation text of the rule to mute while running on the track could be: "When you run on the track and listen to music with headphones, non-private message notifications on your phone will be automatically muted."

[0092] After obtaining the device response information through step 102 above, you can continue to perform the following step 104.

[0093] Step 104: Display the device's response information to the user.

[0094] After displaying the device response information to the user using step 104 above, when the device response information is an execution rule created by the mobile terminal based on the natural language text input by the user, in order to enable the user to modify the generated execution rule or to explain the generated execution rule to the user, the method may further perform the following steps (1) to (5):

[0095] (1) Obtain the user's natural language information, perform semantic understanding on the natural language information, and obtain the semantic understanding result of the natural language information;

[0096] (2) When the semantic understanding result indicates that the user needs to modify the created execution rule, the large language model is used to process the markup language file of the execution rule, the first perceptron list, the second perceptron list, and the natural language information to obtain the modified execution rule;

[0097] (3) When the semantic understanding result indicates that the user is querying the natural language explanation text of the execution rule, obtain the execution rule that needs to form the natural language explanation text;

[0098] (4) Using a large language model, the execution rules that need to form natural language explanation text and the natural language information are processed to obtain the natural language explanation text of the execution rules queried by the user.

[0099] (5) Display the natural language explanation text of the execution rules requested by the user to the user.

[0100] In step (1) above, the natural language information is semantically understood using a large language model to obtain the semantic understanding result of the natural language information. The specific process is existing technology and will not be described in detail here.

[0101] The semantic understanding results include: modified execution rules and natural language explanations of user queries for execution rules.

[0102] In step (2) above, the specific process of using a large language model to process the markup language file of the execution rule, the first perceptron list, the second perceptron list, and natural language information to obtain the modified execution rule is the prior art.

[0103] In step (3) above, if the semantic understanding result is a natural language interpretation text for a user query to execute the rule, the semantic understanding result carries a rule identifier for the execution rule that needs to form the natural language interpretation text. Using the rule identifier in the semantic understanding result, the execution rule that needs to form the natural language interpretation text can be queried from the already created execution rules.

[0104] In step (4) above, the specific process of using a large language model to process the execution rules that need to form natural language interpretation text and the natural language information to obtain the natural language interpretation text of the execution rules queried by the user is existing technology and will not be described in detail here.

[0105] The execution process of the created execution rules will be explained below:

[0106] During the process of obtaining contextual information through the set sensors, the mobile terminal will use the obtained information to traverse the contextual information in the conditional expression of the set execution rule. When it is confirmed that the obtained contextual information completely matches the contextual information in the conditional expression of the execution rule, the API code text in the execution rule that completely matches the obtained contextual information will be executed, and the mobile terminal will be controlled to complete the action of executing the execution rule that completely matches the obtained contextual information.

[0107] In summary, this embodiment proposes a natural language processing method. Upon receiving natural language text input from a user, it acquires contextual information about the user's environment within a preset timeframe and processes the natural language text and contextual information to obtain device response information for the user's input. Compared to related technologies where incomplete user input prevents accurate understanding of the user's intent, this method combines the user's input natural language text with contextual information about the user's environment to understand the user's expressed intent. Even if the user's input is incomplete or unclear, the method can more accurately understand the user's intent and respond accordingly. This makes the mobile terminal's response more aligned with the user's intent, improving the user experience.

[0108] Example 2

[0109] The natural language processing device proposed in this embodiment is used to execute the natural language processing method proposed in Embodiment 1 above.

[0110] See Figure 2 The diagram shows the structure of a natural language processing device. This embodiment proposes a natural language processing device, including:

[0111] The acquisition module 200 is used to acquire contextual information of the user's current situation within a preset time period when it receives natural language text input by the user. The contextual information includes: time information, location information, the application currently being used by the user, and the user's motion status information.

[0112] Processing module 202 is used to process the natural language text and the contextual information to obtain device response information for the natural language text input by the user;

[0113] The display module 204 is used to display the device's response information to the user.

[0114] Specifically, the processing module 202 is used for:

[0115] Obtain multiple pre-set API code texts, each of which includes: the API's identifier, definition, parameters used, and implemented function;

[0116] The API code text and the natural language text are processed using a large language model to determine the API code text involved in the natural language text;

[0117] The natural language text and the contextual information are processed to obtain a list of conditional expressions for key phrases in the contextual information;

[0118] The API code text involved in the determined natural language text and the list of conditional expressions for key phrases in the context information are merged to obtain a markup language file;

[0119] The obtained markup language file is used as the device response information for the natural language text input by the user, thereby obtaining the device response information for the natural language text input by the user.

[0120] In summary, this embodiment proposes a natural language processing device that, upon receiving natural language text input by a user, acquires contextual information about the user's situation within a preset time period, and processes the natural language text and contextual information to obtain device response information for the user's input. Compared to related technologies where incomplete user input prevents accurate understanding of the user's intent, this device combines the user's input natural language text with contextual information about the user's situation to understand the user's expressed intent. Even if the user's input is incomplete or unclear, the device can more accurately understand the user's expressed intent and respond accordingly. This makes the mobile terminal's response to the user more aligned with the user's intent, thus improving the user experience.

[0121] Example 3

[0122] This embodiment proposes a computer-readable storage medium storing a computer program. When the computer program is run by a processor, it executes the steps of the natural language processing method described in Embodiment 1 above. For specific implementation details, please refer to Method Embodiment 1, which will not be repeated here.

[0123] In addition, see Figure 3 The diagram shows the structure of an electronic device. This embodiment also proposes an electronic device, which includes a bus 51, a processor 52, a transceiver 53, a bus interface 54, a memory 55, and a user interface 56. The electronic device includes a memory 55.

[0124] In this embodiment, the electronic device further includes: one or more programs stored in the memory 55 and executable on the processor 52, configured to be executed by the processor to perform the one or more programs for the following steps (1) to (3):

[0125] (1) When the mobile terminal receives natural language text input by the user, it obtains context information of the user's current situation within a preset time period, wherein the context information includes: time information, location information, the application currently used by the user, and the user's motion status information;

[0126] (2) Process the natural language text and the context information to obtain device response information for the natural language text input by the user;

[0127] (3) Display the device response information to the user.

[0128] Transceiver 53 is used to receive and send data under the control of processor 52.

[0129] The bus architecture (represented by bus 51) can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 52 and memory represented by memory 55. Bus 51 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be further described in this embodiment. Bus interface 54 provides an interface between bus 51 and transceiver 53. Transceiver 53 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. For example, transceiver 53 receives external data from other devices. Transceiver 53 is used to transmit data processed by processor 52 to other devices. Depending on the nature of the computing system, a user interface 56 may also be provided, such as a keypad, display, speaker, microphone, or joystick.

[0130] Processor 52 is responsible for managing bus 51 and general processing, such as running general-purpose operating system 551 as described above. Memory 55 can be used to store data used by processor 52 during operation.

[0131] Optionally, the processor 52 may be, but is not limited to, a central processing unit, a microcontroller, a microprocessor, or a programmable logic device.

[0132] It is understood that the memory 55 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate Synchronous DRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 55 of the systems and methods described in this embodiment is intended to include, but is not limited to, these and any other suitable types of memory.

[0133] In some implementations, memory 55 stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof: operating system 551 and application programs 552.

[0134] The operating system 551 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 552 includes various applications, such as a media player and a browser, used to implement various application functions. Programs implementing the methods of the embodiments of this application can be included in the application program 552.

[0135] In summary, this embodiment proposes a computer-readable storage medium and an electronic device. Upon receiving natural language text input from a user, it acquires contextual information about the user's environment within a preset timeframe and processes the natural language text and contextual information to obtain device response information for the user's input. Compared to related technologies where incomplete user input prevents accurate understanding of the user's intent, this embodiment combines the user's input natural language text with contextual information about the user's environment to understand the user's expressed intent. Even if the user's input is incomplete or unclear, the device can more accurately understand the user's intent and respond accordingly. This makes the mobile terminal's response more aligned with the user's intent, improving the user experience.

[0136] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A natural language processing method, characterized in that, include: When the mobile terminal receives natural language text input by the user, it obtains contextual information about the user's current situation within a preset time period. The contextual information includes: time information, location information, the application currently being used by the user, and the user's motion status information. The natural language text and the contextual information are processed to obtain device response information for the natural language text input by the user; Display the device's response information to the user; The natural language text and the contextual information are processed to obtain device response information for the natural language text input by the user, including: Obtain multiple pre-set API code texts, each of which includes: the API's identifier, definition, parameters used, and implemented function; The API code text and the natural language text are processed using a large language model to determine the API code text involved in the natural language text; The natural language text and the contextual information are processed to obtain a list of conditional expressions for key phrases in the contextual information; The API code text involved in the determined natural language text and the list of conditional expressions for key phrases in the context information are merged to obtain a markup language file; The obtained markup language file is used as the device response information for the natural language text input by the user, thereby obtaining the device response information for the natural language text input by the user.

2. The method according to claim 1, characterized in that, The process of processing the natural language text and the contextual information to obtain a list of conditional expressions for key phrases in the contextual information includes: The natural language text and the contextual information are processed using a large language model to obtain contextual description text; The natural language text, the contextual information, and the contextual description text are processed using a large language model to obtain the first keyword in the contextual description text; A first sensor list is determined to identify the sensors that acquire the context information; the sensors represent devices and application software in the mobile terminal that are capable of obtaining context information about the user's current context. Obtain a second perceptron list containing all perceptrons, and use a large language model to process the first perceptron list, the second perceptron list, and the first keyword to obtain the perceptron corresponding to the first keyword. The perceptron corresponding to the first keyword is generalized to obtain a list of first conditional expressions for the first keyword. The natural language text input by the user is processed to obtain a list of second conditional expressions for the second keyword in the natural language text; The first and second conditional expression lists are merged to obtain a conditional expression list for the key phrases in the context information.

3. The method according to claim 2, characterized in that, The process of processing the natural language text input by the user to obtain a list of second conditional expressions for the second keyword in the natural language text includes: The natural language text and contextual information are processed using a large language model to obtain the second keyword in the contextual description text; The first perceptron list, the second perceptron list, and the second keyword are processed using a large language model to obtain the perceptron corresponding to the second keyword. The perceptron corresponding to the second keyword is generalized to obtain a list of second conditional expressions for the second keyword.

4. The method according to claim 2, characterized in that, When the device response information is an execution rule created by the mobile terminal based on the natural language text input by the user, the process of processing the natural language text and the context information to obtain the device response information of the natural language text input by the user further includes: The markup language file is processed using a large language model to obtain the natural language interpretation text of the execution rule.

5. The method according to claim 4, characterized in that, When the device responds with an execution rule created by the mobile terminal based on natural language text input by the user, the method further includes: Obtain the user's natural language information, perform semantic understanding on the natural language information, and obtain the semantic understanding result of the natural language information; When the semantic understanding result indicates that the user needs to modify the created execution rule, the large language model is used to process the markup language file of the execution rule, the first perceptron list, the second perceptron list, and the natural language information to obtain the modified execution rule. When the semantic understanding result indicates that the user is querying the natural language explanation text of the execution rule, the execution rule that needs to be formed into the natural language explanation text is obtained; Using a large language model, the execution rules that need to be formed into natural language explanation text and the natural language information are processed to obtain the natural language explanation text of the execution rules queried by the user. The system displays a natural language explanation of the execution rules requested by the user.

6. A natural language processing device, characterized in that, include: The acquisition module is used to acquire contextual information of the user's current situation within a preset time period when it receives natural language text input by the user. The contextual information includes: time information, location information, the application currently being used by the user, and the user's motion status information. The processing module is used to process the natural language text and the contextual information to obtain device response information for the natural language text input by the user; The display module is used to show the device's response information to the user; The processing module is specifically used for: Obtain multiple pre-set API code texts, each of which includes: the API's identifier, definition, parameters used, and implemented function; The API code text and the natural language text are processed using a large language model to determine the API code text involved in the natural language text; The natural language text and the contextual information are processed to obtain a list of conditional expressions for key phrases in the contextual information; The API code text involved in the determined natural language text and the list of conditional expressions for key phrases in the context information are merged to obtain a markup language file; The obtained markup language file is used as the device response information for the natural language text input by the user, thereby obtaining the device response information for the natural language text input by the user.

7. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program, when run by a processor, performs the steps of the method described in any one of claims 1-5.

8. An electronic device, characterized in that, The electronic device includes a memory, a processor, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor of the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Using context information to facilitate processing of commands in a virtual assistant

    CN103226949A