Processing method and related device
By determining the instruction group in the voice intelligent interactive system and disassembling based on the literal expression of the user input text or the system capabilities of the device, the problem of low generalization of the existing system is solved, and more efficient processing and stronger adaptability are achieved.
Patent Information
- Application Number
- CN202311462024.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-03
- Publication Date
- 2025-05-06
AI Technical Summary
When processing complex user instructions, existing voice intelligent interaction systems need to be based on designed templates and multiple matching models, resulting in low generalization and suitable for customized development, with more manual customization.
By obtaining the user input text, at least one instruction group is determined, and processed based on the decomposition granularity of the indication, including disassembly based on the literal expression of the user input text or the system capabilities of the device, forming an instruction group of conditional instructions and action instructions.
It improves processing efficiency, enhances generalization, reduces the need for manual customization, and enables the system to better understand and execute users' complex instructions.
Smart Images

Figure CN119945815A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of electronic equipment, and in particular to a processing method and related devices. Background Art
[0002] Artificial Intelligence (AI) is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. Artificial Intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can respond in a similar way to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing and expert systems.
[0003] With the development of artificial intelligence, voice intelligent interaction systems are also more widely used in people's daily lives, such as voice interactive TVs, car voice interactive systems, etc. People can free their hands through voice. In voice command interaction, whether the voice intelligent interaction system can correctly understand the user's language input is an important external manifestation of its intelligence level.
[0004] At present, when controlling equipment / calling tools based on user voice understanding of complex user commands, the natural language commands are first parsed into a set of annotated words through the Chinese word segmentation system based on the user input commands and relational connectives as templates, and then through the syntax tree, using the designed templates and matching the corresponding multiple models multiple times, the user input commands are parsed into a set of atomic commands that can be directly executed by the robot. However, since this method needs to be based on a designed template and processed by matching multiple models multiple times, it is only suitable for customized development and has low generalization. Summary of the invention
[0005] The embodiments of the present application provide a processing method and related devices, which can improve processing efficiency, have strong generalization and less manual customization.
[0006] In a first aspect, an embodiment of the present application provides a processing method, comprising: a device obtains a user input text. Then, the device determines at least one instruction group. The determination of the at least one instruction group is related to a first text and the user input text. The first text indicates the granularity of instruction disassembly. Then, processing is performed based on the at least one instruction group to respond to the user input text.
[0007] In the embodiment of the present application, the determination of at least one instruction group is related to the first text and the user input text. The first text indicates the granularity of instruction disassembly. Then, processing is performed based on at least one instruction group to respond to the user input text. With this method, the instruction group obtained based on the indicated disassembly granularity is more efficient, more generalizable, and less manually customized.
[0008] In a possible implementation, the granularity of the instruction decomposition may be a granularity based on the literal expression of the user input text.
[0009] The granularity of the decomposition based on the literal expression of the user input text can be understood as decomposition based on decomposition habits, general common sense, etc. Decomposition based on the literal expression of the user input text can make the command decomposition result more consistent with the user's original expression, decouple the decomposition model from the system's own capabilities, and support the scenario of one understanding engine connecting to multiple devices with different capabilities.
[0010] In another possible implementation, the granularity of the instruction decomposition may be based on the system capability of the device and the literal expression of the user input text.
[0011] The system capabilities of the device can include the capabilities that the system explicitly supports and the potential capabilities of the system. All these achievable disassembly granularities belong to the system capabilities of the device.
[0012] Combining the literal expression of the user input text and the system capabilities of the device for comprehensive decomposition can make the command decomposition result more consistent with the processing capabilities of the device itself, simplify the complexity of subsequent understanding of the system, reduce clarifications and follow-up questions, and more quickly meet user requests based on device capabilities.
[0013] In another possible implementation, each of the at least one instruction group includes at least one conditional instruction and at least one action instruction, and the determination of the at least one instruction group is related to the second text, the first text, and the user input text. The second text indicates at least one of the relationship between the conditional instruction and the action instruction and the relationship between the action instructions in each instruction group.
[0014] By breaking down the user input text into instruction groups including conditional instructions and action instructions, different categories of instructions can be distinguished, so that conditions and actions are identified using independent modules, reducing the design difficulty of the corresponding modules, supporting more complex instruction forms, and improving the processing efficiency of complex instructions.
[0015] In a possible implementation, the relationship between the conditional instruction and the action instruction includes a conditional dependency relationship in which the action instruction is executed when the conditional instruction is met.
[0016] The relationship between the action instructions includes at least one of the following: a serial order relationship in which a second action instruction is executed after a first action instruction is executed, and a multi-instruction dependency relationship in which a next action instruction is executed after at least two action instructions are executed.
[0017] In this example, the user input text is decomposed into instructions based on the second text and the first text indicating the relationship between conditional instructions and action instructions, and the relationship between action instructions. The instructions that satisfy the relationship between instructions indicated by the second text can be classified, so as to ensure the logic between instructions when decomposing the user input text, and keep the instruction execution results consistent with the execution order, execution conditions and execution results expected by the user input text.
[0018] In a possible implementation, at least one conditional instruction and at least one action instruction of each instruction group in the at least one instruction group are obtained by identifying the at least one instruction group. Then, the at least one conditional instruction and the at least one action instruction of each instruction group in the at least one instruction group are spliced to obtain a spliced instruction. The execution result can be obtained by processing based on the spliced instruction.
[0019] In yet another possible implementation, determination of the at least one instruction group is related to a third text, the first text, and the user input text, and the third text indicates a representation form of a relationship between different instructions in the at least one instruction group.
[0020] This example, based on indicating the representation form of the relationship between different instructions, can make the representation form of the instructions obtained by disassembly meet the corresponding requirements, thereby facilitating easy identification during subsequent instruction splicing and execution processing.
[0021] In another possible implementation, the determination of the at least one instruction group is related to the first text, the second text, the third text, and the user input text. That is, the user input text is decomposed into instructions based on the first text, the second text, and the third text to obtain at least one instruction group.
[0022] That is to say, by comprehensively considering the requirements for the granularity of instruction decomposition, the relationship between conditional instructions and action instructions in the instruction group, at least one of the relationships between action instructions, and the requirements for the representation form of the relationship between different instructions, and then performing instruction decomposition on the user input text, a decomposition result that meets the above requirements can be obtained.
[0023] In this example, instruction disassembly is performed based on the first text, the second text, and the third text. In addition to distinguishing conditions and actions, the instruction group obtained by disassembly can also represent the combination and execution order of different actions and different conditions, thereby providing a representation form of instruction disassembly results for more complex scenarios.
[0024] In a possible implementation, the at least one instruction group includes a first conditional instruction and a second conditional instruction, and also includes a first action instruction, a second action instruction, a third action instruction, and a fourth action instruction, wherein the at least one instruction group is expressed in the form of:
[0025] if (first conditional instruction and second conditional instruction) {first action instruction and second action instruction}
[0026] else{third action instruction [fourth action instruction]};
[0027] The representation indicates that when both the first conditional instruction and the second conditional instruction are satisfied, the first action instruction and the second action instruction are executed, and when at least one of the first conditional instruction and the second condition is not satisfied, the third action instruction and the fourth action instruction are executed, wherein the input of the third action instruction depends on the execution of the fourth action instruction.
[0028] It can be understood that "()" is used to determine the scope of the conditional instruction. That is, the instructions in "()" are conditional instructions. "{}" is used to indicate that there are at least two action instructions. Based on {first action instruction and second action instruction}, it can be seen that the first action instruction and the second action instruction are executed in parallel sequence. {third action instruction [fourth action instruction]} means that the input of the third action instruction depends on the execution of the fourth action instruction. In other words, the fourth action instruction is executed first and then the third action instruction is executed.
[0029] In a possible implementation, the first text and the user input text are concatenated to obtain a concatenated text, and then the concatenated text is input into a preset model for processing to obtain the at least one instruction group.
[0030] Among them, the preset model can be a generative model. In this example, at least one instruction group can be obtained based on one model processing. This solution adapts the paradigm of pre-training of the large model itself, converts the semantic parsing task into a natural language generation task, and directly outputs the disassembly result in text form. Processing is based on the first text, which can improve the disassembly efficiency, low latency, and strong generalization.
[0031] In a possible implementation, the method further includes: processing the user input text to obtain the processed user input text. Then, the determination of at least one instruction group is related to the first text and the processed user input text. That is, the instruction is disassembled based on the first text and the processed user input text to obtain the at least one instruction group.
[0032] The processing may include deleting or supplementing words in the user input text, so that the processed user input text is more standardized.
[0033] For example, the colloquial words in the user input text are deleted or omitted. For example, if the user input text is "help me close the window", the processed user input text is "close the window".
[0034] Alternatively, the unclear words in the user input text are supplemented. For example, if the user input text is "help me turn up the air conditioner", the processed user input text is "turn up the air conditioner temperature by 10%". Another example is that the user input text is "turn on the air conditioner and humidifier", and the processed user input text is "turn on the air conditioner and humidifier in the bedroom".
[0035] This example processes the user input text to better adapt to the capabilities of large models, making it simple to adjust the prompt. It is easier to adjust the disassembly granularity, disassembly form, and the instruction dependency order that needs to be supported for different task scenarios. It is more scalable and can effectively improve the efficiency of disassembly. The standardized instructions greatly reduce the ambiguity of understanding caused by the user's spoken expression, reduce the complexity of subsequent understanding modules, and can be connected to behavior and condition recognition modules with lower processing latency, such as the implementation method based on the rule engine.
[0036] In a second aspect, an embodiment of the present application provides a processing device, which may include an acquisition module, a determination module, and a processing module, wherein:
[0037] The acquisition module is used to obtain the user input text;
[0038] A determination module, configured to determine at least one instruction group, wherein the determination of the at least one instruction group is related to a first text and the user input text, wherein the first text indicates a granularity of instruction disassembly;
[0039] A processing module is used to perform processing based on the at least one instruction group to respond to the user input text.
[0040] In a possible implementation, the granularity of the instruction disassembly includes one of the following granularities:
[0041] The granularity of the decomposition based on the literal expression of the user input text, or,
[0042] The granularity of the decomposition is based on the system capabilities of the device and the literal expression of the user input text.
[0043] In one possible implementation, each of the at least one instruction group includes at least one conditional instruction and at least one action instruction, and the determination of the at least one instruction group is related to a second text, the first text, and the user input text, and the second text indicates at least one of the relationship between the conditional instructions and the action instructions and the relationship between the action instructions in each instruction group.
[0044] In a possible implementation, the relationship between the conditional instruction and the action instruction includes a conditional dependency relationship in which the action instruction is executed when the conditional instruction is met;
[0045] The relationship between the action instructions includes at least one of the following:
[0046] A serial order relationship in which a second action instruction is executed after a first action instruction is executed, and a multi-instruction dependency relationship in which the next action instruction is executed after at least two action instructions are executed.
[0047] In a possible implementation, the processing module is used to:
[0048] Performing identification processing on the at least one instruction group to obtain at least one conditional instruction and at least one action instruction of each instruction group in the at least one instruction group;
[0049] splicing the at least one conditional instruction and the at least one action instruction of each instruction group in the at least one instruction group to obtain a spliced instruction;
[0050] Processing is performed based on the splicing instruction to obtain an execution result.
[0051] In a possible implementation, the determination of the at least one instruction group is related to a third text, the first text, and the user input text, and the third text indicates a representation form of a relationship between different instructions in the at least one instruction group.
[0052] In a possible implementation, the at least one instruction group includes a first conditional instruction and a second conditional instruction, and also includes a first action instruction, a second action instruction, a third action instruction, and a fourth action instruction, wherein the at least one instruction group is expressed in the form of:
[0053] if (first conditional instruction and second conditional instruction) {first action instruction and second action instruction}
[0054] else{third action instruction [fourth action instruction]};
[0055] The representation indicates that when both the first conditional instruction and the second conditional instruction are satisfied, the first action instruction and the second action instruction are executed, and when at least one of the first conditional instruction and the second condition is not satisfied, the third action instruction and the fourth action instruction are executed, wherein the input of the third action instruction depends on the execution of the fourth action instruction.
[0056] In a possible implementation manner, the determining module is used to:
[0057] Concatenate the first text and the user input text to obtain a concatenated text;
[0058] The concatenated text is input into a preset model for processing to obtain the at least one instruction group.
[0059] In a third aspect, an embodiment of the present application provides a processing device, comprising at least one processor and a memory, wherein the at least one processor executes a program or instruction stored in the memory so that the processing device implements the method described in any one of the first aspects.
[0060] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium for storing a computer program. When the computer program is executed on a computer or a processor, the computer or the processor implements the method described in any one of the first aspects.
[0061] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed on a computer or a processor, the computer or the processor implements the method described in any one of the first aspects.
[0062] In a sixth aspect, an embodiment of the present application provides a chip, comprising at least one processor and a memory, wherein the at least one processor executes a program or instruction stored in the memory so that the method described in any one of the first aspects is implemented.
[0063] The beneficial effects of the technical solutions provided in the second to sixth aspects of the present application and some implementation methods can refer to the beneficial effects of the technical solution of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The following is a brief introduction to the drawings used in describing the embodiments.
[0065] Figure 1 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application;
[0066] Figure 2It is a schematic diagram of the software structure of an electronic device provided in an embodiment of the present application;
[0067] Figure 3 It is a flowchart of a processing method provided in an embodiment of the present application;
[0068] Figure 4a is a schematic diagram of another processing method provided in an embodiment of the present application;
[0069] Figure 4b This is a schematic diagram of an interface provided by an embodiment of the present application;
[0070] Figure 5 It is a structural schematic diagram of a processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0071] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the embodiments of the present application.
[0072] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0073] The terms "first" and "second" and the like in the description and drawings of the embodiments of the present application are used to distinguish different objects, or to distinguish different processing of the same object, rather than to describe a specific order of objects.
[0074] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the embodiments of the present application are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices.
[0075] It should be noted that in the description of the embodiments of the present application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as having priority or advantage over other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present related concepts in a specific way.
[0076] First, the electronic device involved in the embodiment of the present application is introduced. Figure 1 , Figure 1 It is a schematic diagram of the structure of the electronic device 100 provided in an embodiment of the present application.
[0077] The electronic device 100 may be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, an in-vehicle device, a smart home device and / or a smart city device. The embodiments of the present application do not impose any special restrictions on the specific type of the electronic device.
[0078] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0079] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0080] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0081] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.
[0082] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or cyclically used. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0083] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0084] It is understandable that the interface connection relationship between the modules illustrated in the embodiment of the present invention is only a schematic illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0085] The charging management module 140 is used to receive charging input from a charger. The charger may be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 may receive charging input from a wired charger through the USB interface 130. In some wireless charging embodiments, the charging management module 140 may receive wireless charging input through a wireless charging coil of the electronic device 100. While the charging management module 140 is charging the battery 142, it may also power the electronic device through the power management module 141.
[0086] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and supplies power to the processor 110, the internal memory 121, the display screen 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle number, battery health status (leakage, impedance), etc. In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.
[0087] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0088] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of antennas. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0089] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc., applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0090] The modem processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be sent into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After the low-frequency baseband signal is processed by the baseband processor, it is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to a speaker 170A, a receiver 170B, etc.), or displays an image or video through a display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0091] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, modulates the frequency of the electromagnetic wave signal and performs filtering, and sends the processed signal to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, modulate the frequency of it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0092] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0093] The electronic device 100 implements the display function through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.
[0094] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.
[0095] The electronic device 100 can realize the shooting function through ISP, camera 193, video codec, GPU, display screen 194 and application processor.
[0096] ISP is used to process the data fed back by camera 193. For example, when taking a photo, the shutter is opened, and the light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to ISP for processing and converts it into an image visible to the naked eye. ISP can also perform algorithm optimization on the noise and brightness of the image. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, ISP can be set in camera 193.
[0097] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0098] The digital signal processor is used to process digital signals, and can process not only digital image signals but also other digital signals. For example, when the electronic device 100 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0099] Video codecs are used to compress or decompress digital videos. The electronic device 100 may support one or more video codecs. Thus, the electronic device 100 may play or record videos in a variety of coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0100] NPU is a neural network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, it can quickly process input information and can also continuously self-learn. Through NPU, applications such as intelligent cognition of electronic device 100 can be realized, such as image recognition, face recognition, voice recognition, text understanding, etc.
[0101] The internal memory 121 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM).
[0102] Random access memory may include static random-access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM, for example, the fifth generation DDR SDRAM is generally referred to as DDR5 SDRAM), etc.;
[0103] Non-volatile memory may include disk storage devices and flash memory.
[0104] Flash memory can be divided into NOR FLASH, NAND FLASH, 3D NAND FLASH, etc. according to the operating principle; can be divided into single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc. according to the storage unit potential level; can be divided into universal flash storage (UFS), embedded multi media card (eMMC), etc. according to the storage specification.
[0105] The random access memory can be directly read and written by the processor 110, and can be used to store executable programs (such as machine instructions) of the operating system or other running programs, and can also be used to store user and application data.
[0106] The non-volatile memory may also store executable programs and user and application data, etc., and may be loaded into the random access memory in advance for direct reading and writing by the processor 110 .
[0107] The external memory interface 120 can be used to connect to an external non-volatile memory to expand the storage capacity of the electronic device 100. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, files such as music and videos are stored in the external non-volatile memory.
[0108] The electronic device 100 can implement audio functions such as music playing and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0109] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be arranged in the processor 110, or some functional modules of the audio module 170 can be arranged in the processor 110.
[0110] The speaker 170A, also called a "speaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0111] The receiver 170B, also called a "earpiece", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or voice message, the voice can be received by placing the receiver 170B close to the human ear.
[0112] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to microphone 170C to input the sound signal into microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, realize directional recording function, etc.
[0113] The earphone interface 170D is used to connect a wired earphone and can be a USB interface 130 or a 3.5 mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0114] The gyro sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 180B. For example: when a user shakes the electronic device, the electronic device determines the intensity of the shake based on the gyro sensor 180B. Optionally, the gyro sensor 180B obtains the angle change of the electronic device over a period of time, and then calculates the angular acceleration to determine the magnitude of the shake intensity. In some embodiments, shaking operations of different intensities can correspond to different operation instructions. For example: when a shaking operation with a shaking intensity not less than a first threshold acts on the electronic device, an instruction for text display is executed.
[0115] The accelerometer 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device, and is applied to applications such as horizontal and vertical screen switching and pedometers. For example: when a user shakes the electronic device, the electronic device determines the intensity of the shake based on the accelerometer 180E. Optionally, the accelerometer 180E determines the magnitude of the shake intensity by obtaining the acceleration values of the electronic device in the x, y, and z axes at that moment. In some embodiments, shaking operations of different intensities may correspond to different operating instructions. For example: when a shaking operation with a shaking intensity not less than a first threshold is applied to the electronic device, an instruction to display text is executed.
[0116] The pressure sensor 180A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be set on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can be a parallel plate including at least two conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on the display screen 194, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. The air pressure sensor 180C is used to measure the air pressure. The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case. The distance sensor 180F is used to measure the distance. The electronic device 100 can measure the distance by infrared or laser. The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The proximity light sensor 180G may also be used for automatic unlocking and locking of the screen in the leather case mode and the pocket mode. The ambient light sensor 180L is used to sense the brightness of the ambient light. The fingerprint sensor 180H is used to collect fingerprints. The temperature sensor 180J is used to detect the temperature. The touch sensor 180K, also known as a "touch device". The touch sensor 180K may be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as a "touch screen". The touch sensor 180K is used to detect touch operations acting on or near it. The touch sensor may pass the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation may be provided through the display screen 194. In other embodiments, the touch sensor 180K may also be disposed on the surface of the electronic device 100, which is different from the position where the display screen 194 is located. The bone conduction sensor 180M may obtain a vibration signal.
[0117] The buttons 190 include a power button, a volume button, etc. The motor 191 can generate a vibration prompt. The indicator 192 can be an indicator light, which can be used to indicate the charging status, power change, messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect a SIM card.
[0118] The software system of the electronic device 100 may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present invention, the Android system of the layered architecture is taken as an example to exemplify the software structure of the electronic device 100.
[0119] Figure 2 1 is a software structure block diagram of the electronic device 100 according to an embodiment of the present invention.
[0120] The layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system library, and the kernel layer.
[0121] The application layer can include a series of application packages.
[0122] like Figure 2 As shown, the application package may include camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message and other applications.
[0123] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0124] like Figure 2 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0125] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0126] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0127] The view system includes visual controls, such as controls for displaying text, controls for displaying images, etc. The view system can be used to build applications. A display interface can be composed of one or more views. For example, a display interface including a text notification icon can include a view for displaying text and a view for displaying images.
[0128] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including connecting, hanging up, etc.).
[0129] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0130] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages and can disappear automatically after a short stay without user interaction. For example, the notification manager is used to notify download completion, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as notifications of applications running in the background, or a notification that appears on the screen in the form of a dialog window. For example, a text message is displayed in the status bar, a prompt sound is emitted, an electronic device vibrates, an indicator light flashes, etc.
[0131] Android Runtime includes core libraries and virtual machines. Android runtime is responsible for scheduling and management of the Android system.
[0132] The core library consists of two parts: one part is the function that needs to be called by the Java language, and the other part is the Android core library.
[0133] The application layer and the application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object life cycle management, stack management, thread management, security and exception management, and garbage collection.
[0134] The system library may include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0135] The surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications.
[0136] The media library supports playback and recording of a variety of commonly used audio and video formats, as well as static image files, etc. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0137] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0138] A 2D graphics engine is a drawing engine for 2D drawings.
[0139] The kernel layer is the layer between hardware and software. The kernel layer contains at least display driver, camera driver, audio driver, and sensor driver.
[0140] Hereinafter, the embodiments of the present application will be described in detail with reference to the accompanying drawings.
[0141] Reference Figure 3 As shown in FIG. 1 , it is a flow chart of a processing method provided in an embodiment of the present application. Figure 3 The processing method shown may include steps 301-303. It should be understood that for the convenience of description, this application is described in the order of 301-303, and is not intended to limit the execution to the above order. The embodiment of the present application does not limit the execution order, execution time, execution number, etc. of the above one or more steps. The following description takes the execution subject of steps 301-303 of the processing method as an example of a device, and this application is also applicable to other execution subjects. Steps 301-303 are as follows:
[0142] 301. The device obtains user input text.
[0143] The device may directly obtain the user input text based on the user input, or may obtain the user input text from other devices, etc. This solution does not impose any restrictions on this.
[0144] Optionally, the device may also obtain the above-mentioned user input text, etc. by performing text conversion on the user input voice, etc. For example, the device may receive an audio signal through a microphone. Among them, when the user speaks a voice around the device, the audio signal includes the user's voice signal, and the user's voice signal may include a voice instruction for requesting the device to perform an operation corresponding to the voice instruction. After the microphone receives the audio signal, the device can separate the user's voice signal from the audio signal. Then, the device can obtain the corresponding text information by processing the user's voice signal. Of course, other forms of content input by the user can also be processed to obtain the user input text, and this solution does not limit this.
[0145] 302. Determine at least one instruction group, where the determination of the at least one instruction group is related to the first text and the user input text.
[0146] The first text indicates the granularity of instruction disassembly. It can be understood that the granularity of instruction disassembly is the degree of refinement or detail of instruction disassembly.
[0147] Optionally, the instruction decomposition granularity may be based on the literal expression of the user input text. The decomposition granularity based on the literal expression of the user input text may be understood as decomposition based on decomposition habits, general common sense, and the like.
[0148] Decomposing the text based on the literal expression of the user's input can make the command decomposition result more consistent with the user's original expression, decouple the decomposition model from the system's own capabilities, and support scenarios where one understanding engine connects to multiple devices with different capabilities.
[0149] Alternatively, the granularity of instruction decomposition can be based on the system capabilities of the device and the literal expression of the user input text. The system capabilities of the device can include capabilities explicitly supported by the system of the device, and can also include potential capabilities of the system. These achievable decomposition granularities all belong to the system capabilities of the device.
[0150] Combining the literal expression of the user input text and the system capabilities of the device for comprehensive decomposition can make the command decomposition result more consistent with the processing capabilities of the device itself, simplify the complexity of subsequent understanding of the system, reduce clarifications and follow-up questions, and more quickly meet user requests based on device capabilities.
[0151] In this example, at least one instruction group is determined based on the first text indicating the disassembly granularity and the user input text, so that the instruction group meets the requirements of the disassembly granularity, has strong generalization, and requires less manual customization.
[0152] In a possible implementation, the first text and the user input text are concatenated to obtain a concatenated text, and then the concatenated text is input into a preset model for processing to obtain the at least one instruction group.
[0153] The preset model may be a generative model. A generative model is a method of generating data using a probabilistic model. It can regard data generation as a process of extracting samples from a prior distribution, so that new data can be generated instead of just classifying existing data.
[0154] For example, the generative model can be a Generative Adversarial Network (GAN), which consists of a set of generators and discriminators and can be used to generate different types of data such as images, audio, and text. Of course, other models can also be used, and this solution does not limit this.
[0155] In this example, at least one instruction group can be obtained based on one model processing. This solution adapts to the pre-training paradigm of the large model itself, transforms the semantic parsing task into a natural language generation task, and directly outputs the disassembly results in text form. Processing based on the first text can improve the disassembly efficiency, reduce latency, and improve generalization.
[0156] In a possible implementation, the method further includes: processing the user input text to obtain the processed user input text. Then, the determination of at least one instruction group is related to the first text and the processed user input text. That is, the instruction is disassembled based on the first text and the processed user input text to obtain the at least one instruction group.
[0157] The processing may include deleting or supplementing words in the user input text, so that the processed user input text is more standardized.
[0158] Exemplarily, the colloquial words in the user input text are deleted or omitted. For example, if the user input text is "help me close the window", the processed user input text is "close the window".
[0159] Alternatively, the unclear words in the user input text are supplemented. For example, if the user input text is "help me turn up the air conditioner", the processed user input text is "turn up the air conditioner temperature by 10%". Another example is that the user input text is "turn on the air conditioner and humidifier", and the processed user input text is "turn on the air conditioner and humidifier in the bedroom".
[0160] This example processes the user input text to better adapt to the capabilities of large models, making it simple to adjust the prompt. It is easier to adjust the disassembly granularity, disassembly form, and the instruction dependency order that needs to be supported for different task scenarios. It is more scalable and can effectively improve the efficiency of disassembly. The standardized instructions greatly reduce the ambiguity of understanding caused by the user's spoken expression, reduce the complexity of subsequent understanding modules, and can be connected to behavior and condition recognition modules with lower processing latency, such as the implementation method based on the rule engine.
[0161] The above examples are introduced by taking the determination of at least one instruction group being related to the first text and the user input text as an example. In a possible implementation, the determination of at least one instruction group is related to the second text, the first text, and the user input text. For example, the user input text is decomposed into instructions based on the second text and the first text to obtain at least one instruction group.
[0162] The second text indicates at least one of a relationship between conditional instructions and action instructions and a relationship between action instructions in each instruction group.
[0163] The conditional instruction may be understood as an instruction containing words or meanings such as "if", "when", etc. When the assumption or condition corresponding to the conditional instruction is met, the corresponding instruction is executed.
[0164] The action instruction is also an instruction for performing an operation, such as "turn on the TV", "play music" and the like.
[0165] By breaking down the user input text into instruction groups including conditional instructions and action instructions, different categories of instructions can be distinguished, so that conditions and actions are identified using independent modules, reducing the design difficulty of the corresponding modules, supporting more complex instruction forms, and improving the processing efficiency of complex instructions.
[0166] Optionally, the relationship between the conditional instruction and the action instruction includes a conditional dependency relationship in which the action instruction is executed when the conditional instruction is met. That is, the execution of the action instruction needs to be executed when the conditional instruction is met. Therefore, there is a conditional dependency relationship between the conditional instruction and the action instruction. For example, the user inputs the text "If it rains tomorrow morning, the alarm clock will be set one hour earlier, otherwise it will remain unchanged." Then the conditional instruction is "If it rains tomorrow morning", and the action instruction is "The alarm clock will be set one hour earlier." Of course, there can be other relationships between the conditional instruction and the action instruction, and this solution does not limit this.
[0167] Optionally, the relationship between the action instructions can be a serial order relationship in which the second action instruction is executed after the first action instruction is executed. That is, there is a sequence between the two action instructions. For example, the user input text is "Help me raise the room temperature to 25 degrees, maintain it for 10 minutes and then reduce it to 20 degrees." The first action instruction is "Raise the room temperature to 25 degrees", and the second action instruction is "Maintain it for 10 minutes and then reduce it to 20 degrees."
[0168] Alternatively, the relationship between action instructions can be a multi-instruction dependency relationship in which the next action instruction is executed after at least two action instructions are executed. That is, the instruction of the next action instruction depends on the execution of at least two previous action instructions, and they have a sequence. For example, the user input text is "Start the washing machine, dryer and sweeping robot, and send a WeChat to Xiaomei to say that everything is done after all are completed." The at least two action instructions are "Start the washing machine, dryer and sweeping robot", and the next action instruction is "Send a WeChat to Xiaomei to say that everything is done." Of course, there can be other relationships between action instructions and action instructions, and this solution does not limit this.
[0169] In this example, the user input text is decomposed into instructions based on the second text and the first text indicating the relationship between conditional instructions and action instructions, and the relationship between action instructions. The instructions that satisfy the relationship between instructions indicated by the second text can be classified, so as to ensure the logic between instructions when decomposing the user input text, and keep the instruction execution results consistent with the execution order, execution conditions and execution results expected by the user input text.
[0170] In another possible implementation, the determination of the at least one instruction group is related to a third text, the first text, and the user input text. The third text indicates a representation of the relationship between different instructions in the at least one instruction group. That is, the user input text is decomposed into instructions based on the third text and the first text to obtain at least one instruction group.
[0171] Exemplarily, the at least one instruction group includes a first conditional instruction and a second conditional instruction, and also includes a first action instruction, a second action instruction, a third action instruction and a fourth action instruction, wherein the at least one instruction group is expressed in the form of:
[0172] if (first conditional instruction and second conditional instruction) {first action instruction and second action instruction}
[0173] else{third action instruction [fourth action instruction]}.
[0174] The representation indicates that when both the first conditional instruction and the second conditional instruction are satisfied, the first action instruction and the second action instruction are executed, and when at least one of the first conditional instruction and the second condition is not satisfied, the third action instruction and the fourth action instruction are executed, wherein the input of the third action instruction depends on the execution of the fourth action instruction.
[0175] It can be understood that "()" is used to determine the scope of the conditional instruction. That is, the instructions in "()" are conditional instructions. "{}" is used to indicate that there are at least two action instructions. Based on {first action instruction and second action instruction}, it can be seen that the first action instruction and the second action instruction are executed in parallel sequence. {third action instruction [fourth action instruction]} means that the input of the third action instruction depends on the execution of the fourth action instruction. In other words, the fourth action instruction is executed first and then the third action instruction is executed.
[0176] For example, the user inputs the text "If it rains and is windy tomorrow, help me close the windows and turn on the dehumidifier. Otherwise, help me cache the TV series and send it to the computer in the bedroom." The first conditional instruction is "It will rain tomorrow", the second conditional instruction is "It will be windy tomorrow", the first action instruction is "Close the windows", the second action instruction is "Turn on the dehumidifier"; the third action instruction is "Send the file to the computer in the bedroom", and the fourth action instruction is "Cache the TV series". The instruction group obtained by the disassembly can be expressed as:
[0177] if (it will rain and be windy tomorrow) {close the windows and turn on the dehumidifier}
[0178] else{Send file to the bedroom computer [cache TV series]}.
[0179] Of course, other forms of representation may also be used, and this solution does not limit this.
[0180] This example takes the condition "and" represented by "and" as an example. Of course, the condition "or" can be represented by "or", etc. This solution does not limit this.
[0181] This example, based on indicating the representation form of the relationship between different instructions, can make the representation form of the instructions obtained by disassembly meet the corresponding requirements, thereby facilitating easy identification during subsequent instruction splicing and execution processing.
[0182] In another possible implementation, the determination of the at least one instruction group is related to the first text, the second text, the third text, and the user input text. That is, the user input text is decomposed into instructions based on the first text, the second text, and the third text to obtain at least one instruction group.
[0183] That is to say, by comprehensively considering the requirements for the granularity of instruction decomposition, the relationship between conditional instructions and action instructions in the instruction group, at least one of the relationships between action instructions, and the requirements for the representation form of the relationship between different instructions, and then performing instruction decomposition on the user input text, a decomposition result that meets the above requirements can be obtained.
[0184] In this example, instruction disassembly is performed based on the first text, the second text, and the third text. In addition to distinguishing conditions and actions, the instruction group obtained by disassembly can also represent the combination and execution order of different actions and different conditions, thereby providing a representation form of instruction disassembly results for more complex scenarios.
[0185] The above example is explained by taking the example of obtaining an instruction group, wherein, when the user input text contains multiple groups of conditional instructions, multiple instruction groups can be obtained. For example, the user input text is "If it rains tomorrow, I will not go to the park. If Xiao Ming does not come tomorrow, I will go to the movies." Based on the first group of conditional instructions "It will rain tomorrow", the first instruction group is obtained. Based on the second group of conditional instructions "Xiao Ming will not come tomorrow", the second instruction group is obtained. This solution does not impose any restrictions on this.
[0186] 303. Perform processing based on the at least one instruction group to respond to the user input text.
[0187] In a possible implementation, the at least one instruction group is an executable instruction group, and control can be performed by calling a corresponding interface.
[0188] In another possible implementation, combining Figure 4a and Figure 4b As shown, the user voice inputs: When I say I am tired, help me turn on the seat massage and play soothing music. Based on the instruction disassembly of the user query, at least one instruction group is obtained. The instruction group is an ECA group, which includes an event / condition group (for example, the user says I am tired) and an action group (for example, turn on the seat massage, play soothing music). By identifying and processing the at least one instruction group, at least one conditional instruction (for example, tired) and at least one action instruction (for example, turn on the seat massage, play soothing music) of each instruction group in the at least one instruction group are obtained. Then, the at least one conditional instruction and the at least one action instruction of each instruction group in the at least one instruction group are spliced to obtain a spliced instruction. It can be understood that the splicing not only includes the splicing between the conditional instructions and the action instructions within each instruction group, but also includes the splicing between the instruction groups. Then, based on the splicing instruction, processing is performed to obtain an execution result. Specifically, different interfaces are called based on the splicing instruction to implement different controls to obtain an execution result.
[0189] For example, if the user inputs a command to turn on a dehumidifier, the device controls the dehumidifier by calling the control interface of the dehumidifier.
[0190] In the embodiment of the present application, the determination of at least one instruction group is related to the first text and the user input text. The first text indicates the granularity of instruction disassembly. Then, processing is performed based on at least one instruction group to respond to the user input text. With this method, the instruction group obtained based on the indicated disassembly granularity is more efficient, more generalizable, and less manually customized.
[0191] It should be noted that in the various embodiments of the present application, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between the various embodiments are consistent and can be referenced mutually, and the technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.
[0192] The above detailed description of the method of the embodiment of the present application, the following provides a device of the embodiment of the present application. It can be understood that in the various device embodiments of the present application, the division of multiple units or modules is only a logical division according to function, and is not used as a limitation on the specific structure of the device. In a specific implementation, some functional modules may be subdivided into more small functional modules, and some functional modules may also be combined into one functional module, but no matter whether these functional modules are subdivided or combined, the general process performed by the device is the same. For example, some devices contain a receiving unit and a sending unit. In some designs, the sending unit and the receiving unit can also be integrated into a communication unit, which can implement the functions implemented by the receiving unit and the sending unit. Usually, each unit corresponds to its own program code (or program instruction), and when the program code corresponding to each of these units is run on the processor, the unit is controlled by the processing unit to execute the corresponding process to implement the corresponding function.
[0193] The embodiments of the present application also provide an apparatus for implementing any of the above methods. For example, a processing apparatus is provided that includes units (or means) for implementing each step performed by the device in any of the above methods.
[0194] For example, refer to Figure 5 FIG. 1 is a schematic diagram of a processing device provided in an embodiment of the present application. The processing recommendation device is used to implement the aforementioned processing method, for example Figure 3 The processing method shown.
[0195] like Figure 5 As shown, the device may include an acquisition module 501, a determination module 502 and a processing module 503, wherein:
[0196] The acquisition module 501 is used to acquire the user input text;
[0197] A determination module 502 is used to determine at least one instruction group, wherein the determination of the at least one instruction group is related to a first text and the user input text, wherein the first text indicates a granularity of instruction disassembly;
[0198] The processing module 503 is used to perform processing based on the at least one instruction group to respond to the user input text.
[0199] In a possible implementation, the granularity of the instruction disassembly includes one of the following granularities:
[0200] The granularity of the decomposition based on the literal expression of the user input text, or,
[0201] The granularity of the decomposition is based on the system capabilities of the device and the literal expression of the user input text.
[0202] In one possible implementation, each of the at least one instruction group includes at least one conditional instruction and at least one action instruction, and the determination of the at least one instruction group is related to a second text, the first text, and the user input text, and the second text indicates at least one of the relationship between the conditional instructions and the action instructions and the relationship between the action instructions in each instruction group.
[0203] In a possible implementation, the relationship between the conditional instruction and the action instruction includes a conditional dependency relationship in which the action instruction is executed when the conditional instruction is met;
[0204] The relationship between the action instructions includes at least one of the following:
[0205] A serial order relationship in which a second action instruction is executed after a first action instruction is executed, and a multi-instruction dependency relationship in which the next action instruction is executed after at least two action instructions are executed.
[0206] In a possible implementation, the processing module 503 is used to:
[0207] Performing identification processing on the at least one instruction group to obtain at least one conditional instruction and at least one action instruction of each instruction group in the at least one instruction group;
[0208] splicing the at least one conditional instruction and the at least one action instruction of each instruction group in the at least one instruction group to obtain a spliced instruction;
[0209] Processing is performed based on the splicing instruction to obtain an execution result.
[0210] In a possible implementation, the determination of the at least one instruction group is related to a third text, the first text, and the user input text, and the third text indicates a representation form of a relationship between different instructions in the at least one instruction group.
[0211] In a possible implementation, the at least one instruction group includes a first conditional instruction and a second conditional instruction, and also includes a first action instruction, a second action instruction, a third action instruction, and a fourth action instruction, wherein the at least one instruction group is expressed in the form of:
[0212] if (first conditional instruction and second conditional instruction) {first action instruction and second action instruction}
[0213] else{third action instruction [fourth action instruction]};
[0214] The representation indicates that when both the first conditional instruction and the second conditional instruction are satisfied, the first action instruction and the second action instruction are executed, and when at least one of the first conditional instruction and the second condition is not satisfied, the third action instruction and the fourth action instruction are executed, wherein the input of the third action instruction depends on the execution of the fourth action instruction.
[0215] In a possible implementation, the determining module 502 is configured to:
[0216] Concatenate the first text and the user input text to obtain a concatenated text;
[0217] The concatenated text is input into a preset model for processing to obtain the at least one instruction group.
[0218] For an introduction to the above modules, please refer to the Figure 3 The introduction of the method steps in the illustrated embodiment will not be repeated here.
[0219] It should be understood that the division of the units in the above processing device is only a division of logical functions, and in actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. In addition, the units in the processing device can be implemented in the form of a processor calling software; for example, the processing device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of the units of the device, wherein the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory inside the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits, and the functions of some or all of the units can be realized by designing the hardware circuits, and the hardware circuits can be understood as one or more processors; for example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC), and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be realized by a programmable logic device (PLD), taking a field programmable gate array (FPGA) as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of some or all of the above units. All units of the above devices can be realized in the form of a processor calling software, or in the form of hardware circuits, or in part by a processor calling software, and the rest by hardware circuits.
[0220] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is used to implement one or more steps in the processing method in the aforementioned embodiment.
[0221] An embodiment of the present application further provides a computer program product. When the computer program product is run on a computer, the computer is enabled to execute one or more steps of the processing method in the aforementioned embodiment.
[0222] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the specific descriptions of the corresponding steps in the aforementioned method embodiments, and will not be repeated here.
[0223] It should be understood that in the description of the present application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; wherein A and B can be singular or plural. Also, in the description of the present application, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, wherein a, b, c can be single or multiple. In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first", "second", etc. are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art can understand that the words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not limit them to be necessarily different. Meanwhile, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.
[0224] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the division of the unit is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling, direct coupling, or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0225] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0226] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media integrations. The available medium may be a read-only memory (ROM), or a random access memory (RAM), or a magnetic medium, such as a floppy disk, a hard disk, a tape, a disk, or an optical medium, such as a digital versatile disc (DVD), or a semiconductor medium, such as a solid state disk (SSD), etc.
[0227] The above is only a specific implementation of the embodiment of the present application, but the protection scope of the embodiment of the present application is not limited thereto, and any changes or replacements within the technical scope disclosed in the embodiment of the present application should be included in the protection scope of the embodiment of the present application. Therefore, the protection scope of the embodiment of the present application should be based on the protection scope of the claims.
Claims
1. A processing method, characterized in that: include: The device obtains the user input text; Determine at least one instruction group, wherein the determination of the at least one instruction group is related to a first text and the user input text, wherein the first text indicates a granularity of instruction disassembly; Processing is performed based on the at least one instruction set in response to the user input text.
2. The method according to claim 1, characterized in that The granularity of the instruction disassembly includes one of the following granularities: The granularity of the decomposition based on the literal expression of the user input text, or, The granularity of the decomposition is based on the system capabilities of the device and the literal expression of the user input text.
3. The method according to claim 1 or 2, characterized in that: Each of the at least one instruction group includes at least one conditional instruction and at least one action instruction, and the determination of the at least one instruction group is related to a second text, the first text, and the user input text, and the second text indicates at least one of the relationship between the conditional instructions and the action instructions and the relationship between the action instructions in each instruction group.
4. The method according to claim 3, characterized in that The relationship between the conditional instruction and the action instruction includes a conditional dependency relationship in which the action instruction is executed when the conditional instruction is met; The relationship between the action instructions includes at least one of the following: A serial order relationship in which a second action instruction is executed after a first action instruction is executed, and a multi-instruction dependency relationship in which the next action instruction is executed after at least two action instructions are executed.
5. The method according to any one of claims 1 to 4, characterized in that: The processing based on the at least one instruction group to respond to the user input text includes: Performing identification processing on the at least one instruction group to obtain at least one conditional instruction and at least one action instruction of each instruction group in the at least one instruction group; splicing the at least one conditional instruction and the at least one action instruction of each instruction group in the at least one instruction group to obtain a spliced instruction; Processing is performed based on the splicing instruction to obtain an execution result.
6. The method according to any one of claims 1 to 5, characterized in that: The determination of the at least one instruction group is related to a third text, the first text, and the user input text, the third text indicating a representation form of a relationship between different instructions in the at least one instruction group.
7. The method according to claim 6, characterized in that The at least one instruction group includes a first conditional instruction and a second conditional instruction, and also includes a first action instruction, a second action instruction, a third action instruction and a fourth action instruction, wherein the at least one instruction group is expressed in the form of: if (first conditional instruction and second conditional instruction) {first action instruction and second action instruction} else{third action instruction [fourth action instruction]}; The representation indicates that when both the first conditional instruction and the second conditional instruction are satisfied, the first action instruction and the second action instruction are executed, and when at least one of the first conditional instruction and the second condition is not satisfied, the third action instruction and the fourth action instruction are executed, wherein the input of the third action instruction depends on the execution of the fourth action instruction.
8. The method according to any one of claims 1 to 7, characterized in that: The determining of at least one instruction group comprises: Concatenate the first text and the user input text to obtain a concatenated text; The concatenated text is input into a preset model for processing to obtain the at least one instruction group.
9. A processing device, characterized in that: include: The acquisition module is used to obtain the user input text; A determination module, configured to determine at least one instruction group, wherein the determination of the at least one instruction group is related to a first text and the user input text, wherein the first text indicates a granularity of instruction disassembly; A processing module is used to perform processing based on the at least one instruction group to respond to the user input text.
10. The device according to claim 9, characterized in that The granularity of the instruction disassembly includes one of the following granularities: The granularity of the decomposition based on the literal expression of the user input text, or, The granularity of the decomposition is based on the system capabilities of the device and the literal expression of the user input text.
11. The device according to claim 9 or 10, characterized in that Each of the at least one instruction group includes at least one conditional instruction and at least one action instruction, and the determination of the at least one instruction group is related to a second text, the first text, and the user input text, and the second text indicates at least one of the relationship between the conditional instructions and the action instructions and the relationship between the action instructions in each instruction group.
12. The device according to claim 11, characterized in that The relationship between the conditional instruction and the action instruction includes a conditional dependency relationship in which the action instruction is executed when the conditional instruction is met; The relationship between the action instructions includes at least one of the following: A serial order relationship in which a second action instruction is executed after a first action instruction is executed, and a multi-instruction dependency relationship in which the next action instruction is executed after at least two action instructions are executed.
13. The device according to any one of claims 9 to 12, characterized in that The processing module is used to: Performing identification processing on the at least one instruction group to obtain at least one conditional instruction and at least one action instruction of each instruction group in the at least one instruction group; splicing the at least one conditional instruction and the at least one action instruction of each instruction group in the at least one instruction group to obtain a spliced instruction; Processing is performed based on the splicing instruction to obtain an execution result.
14. The device according to any one of claims 9 to 13, characterized in that The determination of the at least one instruction group is related to a third text, the first text, and the user input text, the third text indicating a representation form of a relationship between different instructions in the at least one instruction group.
15. The device according to claim 14, characterized in that The at least one instruction group includes a first conditional instruction and a second conditional instruction, and also includes a first action instruction, a second action instruction, a third action instruction and a fourth action instruction, wherein the at least one instruction group is expressed in the form of: if (first conditional instruction and second conditional instruction) {first action instruction and second action instruction} else{third action instruction [fourth action instruction]}; The representation indicates that when both the first conditional instruction and the second conditional instruction are satisfied, the first action instruction and the second action instruction are executed, and when at least one of the first conditional instruction and the second condition is not satisfied, the third action instruction and the fourth action instruction are executed, wherein the input of the third action instruction depends on the execution of the fourth action instruction.
16. The device according to any one of claims 9 to 15, characterized in that The determining module is used to: Concatenate the first text and the user input text to obtain a concatenated text; The concatenated text is input into a preset model for processing to obtain the at least one instruction group.
17. A processing device, comprising at least one processor and a memory, characterized in that: The at least one processor executes a program or instruction stored in the memory, so that the processing device implements the method according to any one of claims 1 to 8.
18. A computer-readable storage medium for storing a computer program, characterized in that: When the computer program is executed on a computer or a processor, the computer or the processor is enabled to implement the method according to any one of claims 1 to 8.
19. A computer program product, comprising instructions, characterized in that: When the instructions are executed on a computer or a processor, the computer or the processor is enabled to implement the method according to any one of claims 1 to 8.
20. A chip comprising at least one processor and a memory, characterized in that: The at least one processor executes the program or instruction stored in the memory so that the method according to any one of claims 1 to 8 is implemented.