Interaction method and system based on natural language, and storage medium
Through a natural language-based interaction method, combining the user's initial input information and the tag information of the operable object, the target input information and operations are determined, and operation feedback is generated, which solves the problem of high threshold for use of existing software and improves user experience and operation efficiency.
Patent Information
- Application Number
- PCT/CN2024/074502
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2024-01-29
- Publication Date
- 2025-05-30
AI Technical Summary
The threshold for using existing application software is high, making it difficult for users to understand product functions, resulting in the inability to use or full use of the software's functions, especially the poor user experience of novice users.
The natural language-based interaction method is adopted, and the method executed by the processor includes determining the target input information of the user based on the initial input information of the user and the tag information of the operable object; determining the target operation based on the initial input information and/or the target input information, including display screen operation, input error correction operation and intelligent customer service interaction; and generating operation feedback to provide a user modification window.
It simplifies users' operations on software, improves user experience and efficiency, and helps users use software or web pages more simply and clearly.
Smart Images

Figure CN2024074502_30052025_PF_FP_ABST
Abstract
Description
A natural language-based interactive method, system, and storage medium
[0001] Cross-references
[0002] This application claims priority to Chinese application No. 202311558786.4, filed on November 21, 2023, and the entire contents of the above application are incorporated herein by reference. Technical Field
[0003] This specification relates to the field of human-computer interaction technology, and in particular to a natural language-based interaction method, system, and storage medium. Background Art
[0004] Currently, many application software (such as audio / video playback software, audio and video auxiliary software, other audio application software and other application software) have high usage thresholds. Users may not have enough understanding of product functions, or may not be able to use or fully use the software functions. Most novice users lack a good user experience.
[0005] Therefore, it is hoped to provide an interaction method, system and storage medium based on natural language to simplify the user's operation of the software and thus improve the user experience.
[0006] Summary of the Invention
[0007] One of the embodiments of this specification provides an interaction method based on natural language, which is executed by a processor and includes: determining the user's target input information based on the user's initial input information and the marking information of an operable object, wherein the marking information reflects the characteristics of the operable object; determining the target operation based on the initial input information and / or the target input information, wherein the target operation includes at least one of a display screen operation, an input error correction operation, and an intelligent customer service interaction; and generating operation feedback based on the initial input information and / or the target input information, wherein the feedback interface of the operation feedback includes a user modification window.
[0008] One of the embodiments of this specification provides an interaction method system based on natural language, including: a first determination module, used to determine the user's target input information based on the user's initial input information and the marking information of the operable object, wherein the marking information reflects the characteristics of the operable object; a second determination module, used to determine the target operation based on the initial input information and / or the target input information, the target operation including at least one of a display screen operation, an input error correction operation and an intelligent customer service interaction; and a generation module, used to generate operation feedback based on the initial input information and / or the target input information, the feedback interface of the operation feedback including a user modification window.
[0009] One of the embodiments of this specification provides a computer-readable storage medium, which stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes any one of the natural language-based interaction methods described in the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein:
[0011] FIG1 is a schematic diagram of an application scenario of a natural language-based interaction system according to some embodiments of this specification;
[0012] FIG2 is an exemplary flow chart of a natural language-based interaction method according to some embodiments of this specification;
[0013] FIG3 is an exemplary schematic diagram of an operable object according to some embodiments of this specification;
[0014] FIG4 is an exemplary schematic diagram of a display screen operation according to some embodiments of this specification;
[0015] FIG5 is an exemplary schematic diagram of an input error correction operation according to some embodiments of this specification;
[0016] FIG6 is an exemplary schematic diagram of another input error correction operation according to some embodiments of this specification;
[0017] FIG7 is an exemplary schematic diagram of voice feedback according to some embodiments of this specification;
[0018] FIG8 is an exemplary schematic diagram of card feedback according to some embodiments of this specification;
[0019] FIG9 is an exemplary schematic diagram of a user modification window according to some embodiments of the present specification;
[0020] FIG10 is an exemplary schematic diagram of another card feedback according to some embodiments of this specification;
[0021] FIG11 is an exemplary schematic diagram of an intelligent error correction model according to some embodiments of this specification;
[0022] FIG12 is an exemplary schematic diagram of a prediction model according to some embodiments of this specification;
[0023] FIG13 is an exemplary schematic diagram of prediction results according to some embodiments of this specification;
[0024] FIG14 is an exemplary schematic diagram of a prediction recommendation model according to some embodiments of this specification;
[0025] FIG15 is an exemplary module diagram of a natural language-based interaction system according to some embodiments of the present specification. DETAILED DESCRIPTION
[0026] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.
[0027] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0028] As used in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not refer to the singular but also include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0029] Flowcharts are used throughout this specification to illustrate the operations performed by systems according to embodiments of this specification. It should be understood that preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0030] Audio software or other professional software, web pages, etc. are highly professional, and even after simplifying the interface by means of design simplification, function simplification, etc., it is still relatively complex. In view of this, some embodiments of this specification propose a natural language-based interaction method for simplifying user operations. First, the user's initial input information and the marking information of the operable object are obtained to determine the target input information. Based on the initial input information and / or the target input information, the target operation is determined, and operation feedback is generated. This can implement a natural language-based interaction method, which helps users use software or web pages simply and clearly, and improves user efficiency and experience.
[0031] FIG1 is a schematic diagram of an application scenario of a natural language-based interaction system according to some embodiments of this specification.
[0032] As shown in FIG1 , an application scenario 100 of a natural language-based interactive system may include an audio device 110 , an audio source 120 , a user terminal 130 , a processor 140 , a storage device 150 , and a network 160 .
[0033] The audio device 110 refers to a device with a sound source playback function installed in each composite partition (e.g., composite partition 1, composite partition 2, ... composite partition n). For more information about composite partitions, please refer to the relevant description of Figure 2. For example, the audio device 110 may include a speaker 110-1, a music player 110-2 (iPod, MP3 player, etc.), etc. In some embodiments, the audio device 110 may also have a video playback function, for example, the audio device may also be a television 110-3, etc. In some embodiments, the audio device 110 may have one or more speakers. This application does not limit the type of audio device. The audio device 110 may send basic information of the audio device 110 to the processor 140 via the network 160 for subsequent processing. The basic information of the audio device 110 may include a MAC address, account number, device type, device name, etc. In some embodiments, the audio device 110 may also receive and execute operating instructions issued by the processor 140. For example, the audio device 110 may receive a play instruction issued by the processor 140 and play based on the play instruction.
[0034] In some embodiments, the sound source 120 refers to the sound source that the user desires to play. The sound source 120 can be an audio signal interface, such as USB 120-1, SPDIF interface, AUX interface, etc., or any combination thereof. For another example, the sound source 120 can include an audio signal based on a music push playback method, such as AirPlay 120-2, DLNA 120-3, and QPlay, etc., or any combination thereof. This specification does not limit the type of audio signal source. In some embodiments, the processor 140 sends the audio signal to the matching audio device 110 for playback according to the user's configuration.
[0035] A user refers to a user who uses an audio device. For example, in a home setting, a user may be a family member, etc.; in a business setting, a user may be a property manager, etc. The user may issue user instructions through the user terminal 130. For example, a user may use the user terminal 130 to group audio devices, select at least one audio device to play music, and so on. In some embodiments, the user terminal 130 may obtain user input through various means (e.g., voice or text).
[0036] User terminal 130 refers to a terminal device that provides operational and display functions for user interaction. In some embodiments, user terminal 130 may obtain user instructions based on user input or other operations and send the user instructions to storage device 150 and / or processor 140 for storage and / or subsequent processing. User instructions may include grouping instructions, partitioning instructions, etc.
[0037] In some embodiments, the user terminal 130 may include an input device and an output device. Exemplary input devices may include a keyboard, a mouse, a touch screen, a microphone, or any combination thereof. Exemplary output devices may include a display device, a speaker, a printer, a projector, or any combination thereof. Exemplary display devices may include a liquid crystal display (LCD), a light emitting diode (LED)-based display, a flat panel display, a curved display, a television device, a cathode ray tube (CRT), or any combination thereof.
[0038] In some embodiments, the user terminal 130 may include a mobile device 130 - 1 , a tablet computer 130 - 2 , a laptop computer 130 - 3 , etc. or any combination thereof.
[0039] In some embodiments, the user terminal 130 can process information and / or data. For example, the user terminal 130 can process feature information related to the operation instruction in response to the user's operation instruction. Exemplarily, the user terminal 130 can mark the operable object based on the user's operation instruction. In some embodiments, the user terminal 130 can be used to receive and / or display information sent by the processor 140. For example, the user terminal 130 can display the full list information of audio devices obtained from the processor 140 to the user on the interactive interface. In some embodiments, the user terminal 130 can send the operation instruction and / or feature information related to the operation instruction to one or more components in the application scenario 100 of the natural language-based interactive system. For example, the user terminal 130 can send the user's operation instruction for selecting the prediction result and the operation instruction for marking the operable object to the processor 140 for processing, or send it to the storage device 150 for storage. The user's operation instruction may include grouping instructions, partitioning instructions, marking instructions, etc.
[0040] In some embodiments, the interactive interface can be used to interact with the user (such as real-time interaction). For example, the dialog box of the interactive interface can be used to obtain the user's initial input information. In some embodiments, the dialog box can be any shape, such as a rectangle, an ellipse, an irregular shape, etc. In some embodiments, the dialog box can include icons of input tools (such as a keyboard, a mouse, a microphone, a touch screen, etc.), function buttons (such as send, favorite, forward, save, end the conversation, reduce / enlarge the dialog box, set the dialog box style, etc.), conversation subject (such as the user and the intelligent customer service) information, historical conversation content, etc., or any combination thereof.
[0041] The processor 140 can be used to manage data resources and process data and / or information from at least one component involved in the application scenario 100 of the natural language-based interaction system or an external data source (e.g., a cloud data center). The processor 140 can execute program instructions based on this data, information, and / or processing results, thereby performing one or more functions described in this specification.
[0042] In some embodiments, the processor 140 may receive a grouping instruction sent by the user terminal 130 and generate device grouping information based on the grouping instruction. In some embodiments, the processor 140 may receive a play instruction sent by the user terminal 130 and then send the play instruction to the corresponding audio device 110. In some embodiments, the processor 140 may receive a user operation instruction sent by the user terminal 130 and automatically perform audio source pairing and playback based on the user operation instruction.
[0043] In some embodiments, the processor 140 may receive input from the user terminal 130. The user terminal 130 may include a touch screen, and the user may click or drag on the touch screen to input. For another example, the user terminal 130 may include a microphone, and the user may use the microphone to input voice. For another example, the user terminal 130 may include a camera, and the camera may capture user gestures as input. For another example, the user terminal 130 may include an external mouse, and the user may use the mouse to input. For another example, the user terminal 130 may include an external keyboard, and the user may use the keyboard to input text.
[0044] In some embodiments, processor 140 may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-core processing device). As an example only, processor 140 may include a central processing unit (CPU), a graphics processing unit (GPU), or any combination thereof.
[0045] The storage device 150 can be used to store data and / or instructions. For example, the storage device 150 can be used to store user operation instructions transmitted by the user terminal 130 via the network 160. For another example, the storage device 150 can also be used to store one or more instruction data issued by the processor 140 to the user terminal 130 or the audio device 110. In some embodiments, the storage device 150 can also store data reported by the audio device 110, such as basic information reported by the audio device 110. In some embodiments, data communication can be performed between the storage device 150 and the processor 140 via the network 160, and the storage device 150 can also be part of the processor 140.
[0046] In some embodiments, the storage device 150 may include random access memory (RAM), read only memory (ROM), mass storage, the like, or any combination thereof.
[0047] The network 160 can connect the various components of the system and / or connect the system with external resources. The network 160 enables communication between the various components and with other components outside the system, facilitating the exchange of data and / or information.
[0048] In some embodiments, one or more components of the application scenario 100 of the natural language-based interactive system may transmit data to other components of the application scenario 100 of the natural language-based interactive system via the network 160. For example, the processor 140 may obtain information and / or data from the user terminal 130, the audio device 110, or the storage device 150 via the network 160, or may send information and / or data to the user terminal 130 or the storage device 150 via the network 160. In some embodiments, one or more components of the application scenario 100 may also directly communicate data with each other.
[0049] It should be noted that the application scenario 100 of the natural language-based interactive system is provided for illustrative purposes only and is not intended to limit the scope of this application. A person skilled in the art can make various modifications or variations based on the description of this specification. For example, the application scenario 100 of the natural language-based interactive system can implement similar or different functions on other devices. However, such variations and modifications do not deviate from the scope of this application.
[0050] FIG2 is an exemplary flow chart of a natural language-based interaction method according to some embodiments of this specification. In some embodiments, process 200 may be executed by processor 140 or natural language-based interaction system 1500. As shown in FIG2 , process 200 includes the following steps:
[0051] Step 210 : Determine the user's target input information based on the user's initial input information and the tag information of the operable object. In some embodiments, step 210 is executed by the processor 140 or the first determination module 1510 .
[0052] For more information about users, please refer to the relevant description of Figure 1.
[0053] Initial input information refers to the content in the dialog box of the user input interface. The initial input information can include various forms such as text, images, and voice. In some embodiments, the initial input information can include keywords, phrases, or complete sentences input by the user.
[0054] In some embodiments, the processor may obtain user input information and determine it as initial input information. Various input methods include but are not limited to touch input, voice input, image recognition input, and external device input.
[0055] In some embodiments, the user can click an input tool icon on the dialog box to select the corresponding input tool and begin inputting information. For example, after the user clicks the microphone icon, the microphone icon changes to a dynamic form that displays "Recording," allowing the user to enter voice information. In some embodiments, the user can manually click a function button to complete the corresponding operation, or use the aforementioned input tools to complete the corresponding operation by entering the corresponding command. For example, the user can select a certain interactive content (such as a prediction result) sent by the intelligent customer service and click the Save button to add it to the dialog box. For another example, the user can end the current conversation by typing "End Conversation" on the keyboard and sending it to the intelligent customer service. In some embodiments, the dialog box can be automatically invoked or initiated by the user. For example, the dialog box in the interactive interface can automatically activate the dialog box based on the user's usage scenario information to ask the user whether to interact. For example, when the audio system is activated, the user can invoke the dialog box by issuing a voice command to invoke the dialog box, clicking a button to invoke the dialog box, or imitating a gesture to invoke the dialog box.
[0056] Among them, intelligent customer service is a system that achieves natural interaction with users through technologies such as natural language processing and voice recognition.
[0057] An operable object refers to a target that can be operated in the interface, for example, at least one of a device, a composite partition, an audio source, etc. in the interface. Figure 3 is an exemplary schematic diagram of an operable object according to some embodiments of this specification. As shown in Figure 3, the devices included in the operable object may include players, USB devices, etc. Operation refers to clicking, selecting, touching, etc. on the target. A composite partition refers to a virtual area synthesized in the interface, which reflects the real partition of the target area (such as the home space) where the user wants to perform audio-related operations. Audio-related operations may include playing music in a specific partition of the target area, or a specific player. As shown in Figure 3, in the home space, the composite partition may include a kitchen partition, a bedroom partition, etc. The audio source may include music software such as airplay2 and spotify, and may also include other local inputs (such as analog input, etc.).
[0058] It should be noted that the interactive interface includes operable objects and inoperable objects. Inoperable objects refer to objects in the interface that cannot be operated. For example, software feedback information can be about inoperable objects. Software feedback information refers to the relevant information provided by the software in response to the user's initial input information and / or target input information, such as the song being played, device name, partition name, etc.
[0059] Tag information refers to relevant information for marking an operable object. In some embodiments, the tag information may reflect the characteristics of the operable object (such as type, name, etc.). Exemplarily, the tag information is shown in FIG3 , and the composite partition may include default partition 1, study area 2, kitchen 3, and bedroom 4. Default partition 1 can be used to play an audio signal from airplay2; study area 2 is playing an audio signal from airplay2, and the audio signal is Free Summer. Study area 2 includes 4 audio devices, including audio device 1, audio device 2, audio device 3, and audio device 4; kitchen 3 includes audio device 5, and bedroom 4 includes audio device 6; the tag information of mobile storage devices such as audio sources Airable1, airplay2, AUX, roon, Spotify, and U disk are 1, 2, 3, 4, 5, and 6, respectively.
[0060] In some embodiments, the processor can obtain the marking information in a variety of ways. In some embodiments, the processor can obtain the marking information based on the user's operation. The user's operation can include various methods such as manual and input control. In some embodiments, the user can mark the operable objects in a variety of forms, such as Arabic numerals, English letters, representative characters (such as the first letter of the operable object), or personalized settings by the user, etc., and this specification does not limit this. For example, the user can mark the operable object by long pressing it. For another example, the user can input "Mark speaker 1 as 1", and the processor can determine the marking information of speaker 1 as 1. By marking the operable object, the user can simplify the initial input information "Please help me adjust the volume of speaker 1 to 50" to "Please help me adjust the volume of 1 to 50" when inputting later, which effectively simplifies the user's input.
[0061] In some embodiments of this specification, users can refer to corresponding operable objects through marking information. When the name of the operable object is a complex or professional noun, users can quickly control the aforementioned operable object through simple marking information, which is convenient for users to use.
[0062] Target input information refers to the information that the user ultimately sends. Target input information can more completely and accurately reflect the user's intention.
[0063] In some embodiments, the processor may determine the user's target input information in a variety of ways based on the user's initial input information and the tag information of the operable object. For example, the processor may extract keywords based on the initial input information and the tag information, and search the database based on the keywords to determine the target input information. The database may include multiple reference target input information. The database may be constructed based on historical target input information. For example, the keywords extracted by the processor based on the initial input information and the tag information are "1", "volume", and "50". Based on the aforementioned keywords, the processor searches the database and determines that the target input information is "Please help me adjust the volume of speaker 1 to 50".
[0064] In some embodiments, the processor can determine the prediction result through the prediction model based on the initial input information, the currently connected device, and the tag information of the operable object, and determine more content of the target input information. Please refer to the relevant description of Figure 12.
[0065] In some embodiments, the processor may determine the target input information through an input error correction operation based on the initial input information. For more details, please refer to the relevant description after FIG2 .
[0066] In some embodiments, the processor may determine the target input information based on a user's drag instruction on the operation target in the display screen.
[0067] During the user input process, the display screen changes in real time according to the user input.
[0068] Action targets refer to elements on the display screen (such as text, images, links, forms, and other interactive elements).
[0069] The drag command is used to move the operation target on the display screen. In some embodiments, the processor can capture the user's action of moving the operation target using a mouse or touch, etc., and determine the drag command. In some embodiments, the drag command may include the direction of movement, the target movement position, and text information corresponding to the operation target (e.g., Speaker 1, Speaker 2, etc.). The target movement position is a dialog box.
[0070] In some embodiments, the processor can capture the action of the user dragging the operation target into the dialog box by mouse or touch, determine the user's drag instruction, insert the name corresponding to the operation target as part of the target input information into the initial input information, and display it in the dialog box. In some embodiments, the insertion position can be the position of the current insertion point, or the position of the end of the text inserted into the initial input information, or the position of other user-selected insertion points. The insertion point refers to the point in the dialog box used to indicate the text insertion position. In some embodiments, the interactive interface can display the insertion point in the form of a cursor (e.g., a flashing short vertical line, etc.).
[0071] In some embodiments, when entering initial input information, the user can drag the operation targets displayed in the drag interface into the dialog box. Accordingly, the processor can receive the user's drag instruction and convert the operation targets dragged into the dialog box into the corresponding names of the operation targets to speed up user input. For example, in Example 1 of Figure 4, the user can drag the operation targets "Speaker 1" and "Speaker 2" to the dialog box respectively, and the content of the dialog box will change to "Create a partition and specify Speaker 1 and Speaker 2".
[0072] In some embodiments of the present specification, by determining target input information through dragging instructions, the purpose of accelerating user input can be achieved and the user experience can be improved.
[0073] In some embodiments, the operation target may include at least one of a draggable target and a non-draggable target, and the first display modes corresponding to the draggable target and the non-draggable target are different.
[0074] A draggable target is an element in the interface that can be moved.
[0075] Non-draggable targets are elements in the interface that cannot be moved. For example, the corresponding function of an element may be faulty (e.g., the audio file is damaged, cannot be played, cannot be opened, etc.), or the element may be restricted to be immovable, and thus cannot be dragged by the user.
[0076] The first display method is a display method used to distinguish between draggable targets and non-draggable targets. The display method is a way of displaying elements on the display area of the interactive interface. The display method includes but is not limited to: displayed position information, displayed appearance information, hierarchical relationship with other elements in the screen display area, etc. The displayed position information refers to the position of the element in the interactive interface. The displayed appearance information refers to the appearance of the element in the interactive interface, such as font style, layout style, border and background style, animation and transition effects, hierarchical relationship, etc. The hierarchical relationship refers to the hierarchical structure between elements in the interactive interface. When elements are stacked, elements with higher hierarchies are displayed on top, covering elements with lower hierarchies.
[0077] In some embodiments, the first display mode can be implemented in a variety of ways. For example, the draggable and non-draggable targets can be displayed in different colors, different text thicknesses, etc. For example, the draggable target that has been dragged into the dialog box is displayed with a border and gray shadow, the non-draggable target is displayed with a border and black shadow, and the undraggable target is displayed with a border, etc.
[0078] In some embodiments, the processor may preset the display modes corresponding to the draggable object and the non-draggable object as the first display mode. In some embodiments, the processor may also obtain the first display mode through manual input.
[0079] In some embodiments of this specification, different elements (such as draggable, non-draggable, etc.) are presented to users in different ways in the interface, which allows users to quickly select the required draggable targets, increase input convenience, improve input efficiency, and enhance the user experience.
[0080] In some embodiments, the display screen comes from at least one platform, and the at least one platform corresponds to the second display mode of the operation target; and / or the display screen comes from at least one system, and the at least one system corresponds to the third display mode of the operation target, and the second display mode and the third display mode are different.
[0081] System refers to an operating system equipped with a graphical user interface, such as a desktop operating system, a mobile operating system, an embedded operating system, or any combination thereof.
[0082] Since the display screen changes in real time according to the user's input information, the display screen can come from different systems. For example, when the user's initial input information includes operation targets of multiple systems, the operation targets of the display screen can come from different versions of operating systems, such as the previous version of the operating system, the latest version of the operating system, etc.
[0083] Cross-system operation refers to operating (e.g., moving, marking, etc.) draggable objects from different systems on the same display screen.
[0084] A platform is the hardware environment that hosts the operating system. For example, the hardware environment refers to the physical computer system consisting of the processing device and its peripherals, including mobile devices, tablets, desktop computers, etc.
[0085] Since the display screen changes in real time according to the user's input information, the display screen can come from different platforms. For example, when the user's initial input information includes operation targets of multiple platforms, the operation targets of the display screen can come from different platforms, such as Roon and Spotify can come from different platforms.
[0086] Cross-platform refers to operating (e.g., moving, marking, etc.) draggable objects from different platforms on the same display screen.
[0087] The second display mode is used to distinguish the display modes of the operation targets presented on different platforms. In some embodiments, the second display mode includes multiple different display modes, and each platform can correspond to one display mode.
[0088] The third display mode is used to distinguish the display modes of the operation targets presented by different systems. In some embodiments, the third display mode includes multiple different display modes, and each system can correspond to one display mode.
[0089] In some embodiments, the second display mode can be implemented in a variety of ways. For example, different border colors, different background colors, etc. can be used to display the operation targets of different platforms. In some embodiments, the third display mode can be implemented in a variety of ways. For example, different border types can be used to display the operation targets of different systems.
[0090] In some embodiments, the processor may preset display modes corresponding to different platforms as the second display mode. In some embodiments, the processor may also obtain the second display mode through manual input.
[0091] In some embodiments, the processor may preset display modes corresponding to different systems as the third display mode. In some embodiments, the processor may also obtain the third display mode through manual input.
[0092] In some embodiments, the second and / or third display modes are used to display operation targets from different platforms and / or different systems, respectively, so that users can move draggable targets from other platforms and / or systems into a dialog box in the same interactive interface (e.g., an interactive interface of a mobile terminal or a desktop terminal, etc.), thereby achieving cross-platform and cross-system interaction. Mobile terminals refer to related applications used on mobile devices such as mobile phones and tablets. Correspondingly, desktop terminals refer to related applications used on fixed devices such as desktop computers, etc.
[0093] In some embodiments of this specification, cross-platform or cross-system dragging allows users to drag on one platform to control functions on another platform. At the same time, information can be obtained and shared across various devices and systems (e.g., sharing files, music, videos, and other resources), making input more convenient and efficient. Through different display methods, users can quickly distinguish between operation targets from different sources (e.g., from different platforms or different operating systems), improving user convenience and thus improving input efficiency.
[0094] Step 220 , based on the initial input information and / or the target input information, determines the target operation. In some embodiments, step 220 is performed by the processor 140 or the second determination module 1520 .
[0095] The target operation refers to the relevant operation performed by the processor after receiving the initial input information and / or the target input information. In some embodiments, the target operation may include at least one of a display screen operation, an input error correction operation, and intelligent customer service interaction.
[0096] In some embodiments, the processor may analyze and process the initial input information and / or the target input information, and determine the target operation corresponding to the user input in a variety of ways.
[0097] In some embodiments, the processor may trigger a corresponding target operation function based on the user's initial input information and / or target input information.
[0098] For example, in response to the user's initial input information / target input information being "display screen adjustment", "I want to adjust the display screen", etc., the processor may determine that the target operation is a display screen operation and trigger the interactive interface to be in a display screen operation mode.
[0099] For another example, in response to identifying an error in the user's initial input information, the processor can determine that the target operation is an input error correction operation and automatically trigger the error correction function to correct the user input information in real time. For more information about the input error correction operation, please refer to the relevant description below.
[0100] For another example, the processor can determine, based on semantic recognition of the user's initial input and / or target input, that the user wants to inquire about the application's relevant operating tutorials, guidance information, etc., determine the target operation is an intelligent customer service interaction, and trigger the intelligent customer service interaction function. Furthermore, the intelligent customer service interaction function retrieves the corresponding operating tutorials, guidance information, etc. based on the specific information identified by semantic recognition. For more information about intelligent customer service interaction, please refer to the relevant description below.
[0101] Display screen operations refer to operations performed on the display screen of the interactive interface based on the initial input information and / or the target input information. Display screen operations may include switching the display screen (displaying different display screens) based on the initial input information and / or the target input information, scaling, moving, and arranging the operable objects and their information in the interface, and other operations. In some embodiments, display screen operations may also include switching the display screen in real time, scaling, moving, and rearranging the operable objects and their information for display, and the like.
[0102] For example, if the user's initial input information and / or target input information is "Please help me adjust the volume of speaker 2 to 50," the processor can determine, based on the initial input information, that the target operation is to display the screen. The processor can display the screen of speaker 2 in the interface and indicate that the volume of speaker 2 is adjusted to 50.
[0103] For another example, if the user's initial input information and / or target input information is "I want to modify the learning area," the processor may determine, based on the initial input information, that the target operation is a display screen operation. The processor may display only the display screen of the learning area in the interface.
[0104] For another example, the user's initial input information and / or target input information is "Please help me move the learning area upwards by 1 cm", the processor can move the learning area on the interactive interface upwards by 1 cm and display it.
[0105] In some embodiments, the processor may trigger a display screen operation mode after determining that the target operation is a display screen operation. In the display screen operation mode, the user can perform corresponding display screen operations through gestures on the interactive interface. For example, the user can use gestures such as zooming, touching, dragging, and selecting to zoom, move, and arrange the displayed content on the interactive interface.
[0106] As shown in Figure 4, Figure 4 is an exemplary schematic diagram of the display screen operation shown in some embodiments of this specification. In Example 1, the user's initial input information is "Create a partition and specify". The processor can display the speakers that can be specified on the display screen (such as Speaker 1, Speaker 2, Speaker 3 and Speaker 4 in Example 1 in Figure 4) based on the aforementioned initial input information. In Example 2, the user drags Speaker 1 and Speaker 2 into the dialog box based on the display screen operation in Example 1. The processor can insert the names corresponding to Speaker 1 and Speaker 2 to the end of the text of the initial input information in response to the drag instruction, and update the initial input information. The user can continue to input "and play" based on the updated initial input information. The processor can display airplay2, spotify and other operable objects that can be played on the interface based on the initial input information. Among them, the airplay2, spotify, etc. displayed on the interface are streaming media types supported by audio devices such as speaker 1 and speaker 2.
[0107] In some embodiments, the streaming media types supported by the same audio device group can be the same. For example, if an audio device group includes audio devices 1-3, and audio devices 1-3 all support five streaming media types (AirPlay2, Spotify, Roon, DLNA, and Airable), then the streaming media types supported by the audio device group include the five streaming media types (AirPlay2, Spotify, Roon, DLNA, and Airable).
[0108] In some embodiments, the streaming media types supported by the same audio device group can be different. In this case, the supported streaming media types are the union of the streaming media types supported by each audio device in the group. For example, an audio device group includes audio devices 1 to 3. Audio device 1 supports AirPlay 2 and Spotify, audio device 2 supports Spotify and Roon, and audio device 3 supports Airable. In this case, the streaming media types supported by the audio device group are AirPlay 2, Spotify, Roon, and Airable.
[0109] An audio device group refers to one or more audio devices in the initial input information or the target input information.
[0110] In some embodiments, the processor can extract keywords based on the initial input information and / or target input information; determine keyword attributes based on the keywords, and replace the keywords to determine an adjusted keyword group; and search the database based on the adjusted keyword group to determine the content to be displayed on the interface.
[0111] Keywords refer to words related to determining the content to be displayed, such as create and specify in "create a partition and specify".
[0112] Keyword attributes can include operation keywords, name keywords, etc. Operation keywords refer to keywords related to operations, such as create, modify, etc. Name keywords refer to keywords related to the names of operable objects, such as speaker, zone, or tag information reflecting the names of operable objects.
[0113] Replacing a keyword means replacing a name keyword with the original name of the corresponding operable object. For example, if the name keyword "Yang 1" is the identifier for the operable object "Speaker 1," the processor may replace "Yang 1" with "Speaker 1." The adjusted keyword group is a phrase consisting of the replaced keywords and the keywords that do not need to be replaced.
[0114] The database can be built based on historical operation data and operation logic. For example, the keyword "create" can only be used for operable object 1, so the search results for "create" in the database only correspond to operable object 1.
[0115] Exemplarily, the processor may use the adjusted keyword group (create, specify) to search in the database, and determine the search results (speaker 1, speaker 2, speaker 3, speaker 4) as the content displayed on the interface.
[0116] In some embodiments of this specification, the display screen can be changed in real time based on user input through display screen operation, which can provide users with visual prompts and facilitate user operations.
[0117] Input correction refers to the automatic correction of the user's initial input information. The input correction operation can correct the text converted from the user's voice input or other input (such as an image), or it can correct the text entered by the user.
[0118] In some embodiments, the processor can implement input error correction operations in a variety of ways. For example, the processor can perform similarity analysis on the current text and multiple preset texts, and determine the most similar preset text as the corrected text. For another example, the processor can automatically correct the initial input information based on a trained error correction model. The error correction model can be trained based on professional vocabulary, professional phrases, and professional sentence expressions in professional fields such as audio as samples. After training, it can correct professional terms and terminology for audio software, etc.
[0119] In some embodiments, the processor may automatically correct the user's initial input information and determine the corrected initial input information as the target input information.
[0120] For example, the user's initial input information is "turn the volume to 80 and play ox in the living room", the processor can correct the misspelled professional term "ox" to "aux", thereby determining the target input information as "turn the volume to 80 and play aux in the living room".
[0121] As shown in Figure 5, Figure 5 is an exemplary schematic diagram of an input error correction operation according to some embodiments of this specification. The user's initial input information is "Create a partition and specify Speaker 1 and Speaker 2 to enter the partition and play airaplay2." The processor can use the trained error correction model to correct the misspelled professional term "airaplay2" to "airplay2", thereby determining that the target input information is "Create a partition and specify Speaker 1 and Speaker 2 to enter the partition and play airplay2."
[0122] As shown in Figure 6, Figure 6 is an exemplary diagram of another input error correction operation according to some embodiments of this specification. The user's initial input information is "Adjust the volume of all zones to 60 and play music." The processor can use the trained error correction model to correct the incorrectly written professional term "sound" to "volume," thereby determining that the target input information is "Adjust the volume of all zones to 60 and play music."
[0123] In some embodiments of this specification, automatic error correction of initial input information is performed through a trained error correction model, which can improve the accuracy of user input, further shorten the user's trial and error time, and help improve the user's experience.
[0124] In some embodiments, the processor can determine the input information after correction based on the initial input information and / or target input information through an intelligent error correction model and perform automatic error correction. For more details, please refer to the relevant description of Figure 11.
[0125] Intelligent customer service interaction refers to intelligent interaction between a processor and an interactive interface. In some embodiments, intelligent customer service interaction can predict at least one of the following: the user's target input information, the actionable object the user intends to operate, the user's expected usage function, and the question the user intends to ask. For more information on predicting the user's target input information, please refer to the relevant description of Figure 12.
[0126] In some embodiments, intelligent customer service interaction includes predicting the user's emotional state. For more information, please refer to the relevant description of Figure 14.
[0127] In some embodiments, the processor can directly combine the prediction result with the user's current initial input information as the target input information, and extract the keywords of the target input information to predict the operable object or function that the user wants to operate. For example, the user's current initial input information is "speaker 1, speaker 2 then", the processor can combine the prediction result "play" with the initial input information as the target input information "speaker 1, speaker 2 then play", and extract keywords based on the target input information to obtain the content displayed on the interface. For more instructions on extracting keywords and obtaining the content displayed on the interface, please refer to the relevant description below.
[0128] In some embodiments, the intelligent customer service interaction includes predicting the user's estimated usage functions. The processor can determine the user's user characteristics based on the user's behavioral habits; and predict the estimated usage functions based on the user characteristics.
[0129] Estimated usage functions refer to functions that users are expected to use. For example, estimated usage functions may include recommending music, playing songs, opening a specified playlist, etc.
[0130] User behavior refers to the data generated by various user operations on related applications. For example, user behavior may include user browsing data (e.g., viewed songs, videos, etc.), user operation data (e.g., registration, follow, click, close, favorite, comment, feedback, etc.), etc.
[0131] In some embodiments, the processor can obtain user behavior habits in a variety of ways. For example, the processor can extract user behavior habits based on other social networks, software, etc. Exemplarily, the processor can embed relevant interfaces in various social networks, APPs, and websites in audio-related fields. The processor can obtain user behavior habits through relevant interfaces. Exemplarily, the processor obtains the user's operation data at a preset time (such as in the past month) from a third-party music software, and counts the distribution characteristics of the user's operation data (such as the time period for listening to songs, the type of songs, etc.), to determine the user's behavior habits. Among them, the operation data is the data generated by the user's operation on the operation target (such as a function button, a song). For example, the operation can be any operation of the user on the object, such as browsing, clicking, purchasing, commenting, etc. Similarly, the processor can also obtain other user behavior habits based on other social platforms.
[0132] User characteristics refer to objective and / or expressive characteristics that can characterize the user, such as gender, age, income, personality, and behavioral habits.
[0133] In some embodiments, the processor may obtain user characteristics in a common manner, for example, through network transmission, calling an interface, etc.
[0134] In some embodiments, user characteristics may include identity characteristics and / or user operation characteristics. Identity characteristics and user operation characteristics may represent information such as the user's own characteristics, preferences, and personality from different perspectives.
[0135] Identity characteristics refer to characteristics that represent basic information about a user from their perspective. Examples include gender, age, personality, education, income, occupation, risk tolerance, and decision-making preferences.
[0136] User operation characteristics are features that represent various user behavior information. For example, user operation characteristics can include usage characteristics, browsing characteristics, information input characteristics, device usage characteristics, content consumption characteristics, social characteristics, search characteristics, feedback and interaction characteristics, or any combination thereof. Usage characteristics can include commonly used functions and operation characteristics (single-handed vs. two-handed). Browsing characteristics can include characteristics of users browsing or searching for operation targets (e.g., browsing from left to right, top to bottom, or following a specific information acquisition path). Information input characteristics can refer to characteristics of users inputting information. For example, information input characteristics can include one or any combination of input methods such as text, voice, and images. Device usage characteristics include characteristics of users using related applications on different devices (e.g., mobile phones, tablets, and computers), including device usage frequency and duration. Content consumption characteristics can include user preferences for information content (e.g., operation targets or functions), reading characteristics, and interaction methods, such as preference for images, text, video, or audio, or demand for different types of content. Social characteristics refer to user habits on social media, such as the frequency and preferences of interactive behaviors such as posting, forwarding, and commenting. Search characteristics refer to the characteristics of users when searching for information, such as commonly used search engines and search keyword choices. Feedback and interaction characteristics can include user preferences for providing feedback and suggestions, as well as how they resolve problems encountered while using a product or service.
[0137] In some embodiments, the processor can determine the user characteristics of the user based on the user's behavioral habits in a variety of ways. For example, the processor can embed relevant interfaces in various social networks, audio-related APPs, and websites. The processor can obtain the user's identity characteristics through relevant interfaces. The account information may include the user's registration information and the user's operation data on the account. For another example, the processor can perform statistical analysis based on the user's behavioral habits within a preset time period (e.g., the past month, the past six months) to determine the user's operation characteristics.
[0138] In some embodiments, the processor can predict the estimated usage function based on user features in a variety of ways. For example, the predicted estimated usage function can be determined based on a preset table or vector database constructed based on historical data. The preset table / vector database can be a table or database that characterizes the correspondence between user features and estimated usage functions. In some embodiments, the processor can also predict the estimated usage function through a predictive function model. The predictive function model can be a machine learning model such as a trained neural network. The input of the predictive function model can include user features, and the output can include predicted estimated usage functions. The predictive function model can be trained in various feasible ways based on a large number of first training samples with a first label. For example, parameter updates can be performed based on the gradient descent method. The first training sample may include sample user features, which can be obtained based on historical data, and the first label may be the function actually chosen by the sample user to use, which can be obtained by manual or processor annotation. The training method of the predictive function model is similar to the training method of the prediction model, and reference can be made to the relevant description of FIG12.
[0139] In some embodiments of this specification, by collecting statistics on user behavior habits, it is possible to accurately determine the recommended functions that users may want to know about, thereby providing users with better proactive services. User characteristics can more directly reflect user habits and preferences, thereby further increasing the accuracy of predicted estimated usage functions.
[0140] In some embodiments, the processor may determine the question the user intends to ask based on historical user input sequences, user operation sequences, currently connected devices and operable objects, and their tag information in various ways. For example, this may be determined through preset rules or through a vector database. For another example, the processor may use a question prediction model to predict the question the user intends to ask based on historical user input sequences, user operation sequences, currently connected devices and operable objects, and their tag information.
[0141] A historical user input sequence refers to a sequence consisting of the user's initial input information and / or target input information within a preset time, and the preset time can be determined based on actual conditions. A user operation sequence refers to a sequence consisting of the user's operations in the interface before starting to input the dialog box. For more information about operable objects and their marking information, please refer to the relevant description above. Among them, the currently connected device refers to the related device connected to the current interface, such as car Bluetooth, speakers, etc. In some embodiments, the processor can obtain the currently connected device from the interface based on various methods such as wired or wireless connection.
[0142] A question prediction model is a model used to predict questions a user wants to ask. The question prediction model can be a machine learning model. The question prediction model can be trained based on a large number of second training samples with second labels. The training process for the question prediction model is similar to that for the prediction model; see Figure 12 for the relevant description. The second training samples can include sample historical user input sequences, sample user operation sequences, sample currently connected devices, and sample operable objects and their labeling information. The second training samples can be obtained based on historical data. The second label can be the actual question corresponding to the second training sample.
[0143] Some embodiments of this specification can simplify user input and improve user operation efficiency by predicting the questions the user wants to ask. For example, if the processor obtains that the user has been operating the settings of Speaker 1 for the past ten minutes, when the user opens a dialog box, the processor can display the predicted question the user wants to ask (such as "How do I use a certain function of Speaker 1") in the interface, and the user can simplify input by selecting the question in the interface.
[0144] In some embodiments, intelligent customer service interaction can also include automatically generating a target operation video based on initial input information and / or target input information. A target operation video is an operation video used to guide the user to complete a corresponding operation. Target operation videos can include, for example, robot simulation operation videos. For example, if the user's target input information is "How to combine Speaker 1 and Speaker 2 to play a sound source using the combined Speakers 1 and 2," the processor can generate a target operation video corresponding to the question and display it to the user through the interface.
[0145] In some embodiments, the processor can generate the target operation video in a variety of ways. For example, the processor can extract keywords from the initial input information and / or the target input information, and obtain the target operation video from the storage device based on the keywords.
[0146] In some embodiments, the processor may extract keywords from the initial input information and / or the target input information, and directly retrieve existing operation videos from other platforms.
[0147] In some embodiments, the processor can obtain a plurality of pre-set basic operation pictures, operation video clips, complete operation videos and characteristic materials of different devices (such as device models, logos, etc.). Furthermore, the processor can synthesize the aforementioned basic operation pictures, operation video clips, and characteristic materials based on the keywords, device models, etc. in the initial input information and / or the target input information to generate a corresponding target operation video. Alternatively, the processor intercepts the corresponding operation video segment from the aforementioned complete operation video based on the keywords, device models, etc. in the initial input information and / or the target input information, and combines it with the characteristic materials to generate the target operation. The target operation video can be marked with the model and logo information corresponding to the device, without the need to record a video for each device.
[0148] In some embodiments, the processor may generate a target operation video using a video generation model based on the initial input information and / or the target input information.
[0149] The video generation model may be a machine learning model. In some embodiments, the video generation model may include a problem determination layer and a video generation layer.
[0150] The input to the question determination layer may include initial input information and / or target input information, and the output may be the corresponding question. The initial input information and / or target input information input to the video generation model may include the user's query content. The question determination layer may be a natural language processing (NLP) model, for example. The question determination layer may be trained using a large number of third training samples with third labels. The training process for the question determination layer is similar to that of the prediction model; please refer to the relevant description in Figure 12.
[0151] The third training sample can be sample initial input information and / or sample target input information, and can be obtained based on historical data. The third label can be an actual problem corresponding to the third training sample, and can be manually labeled.
[0152] The video generation layer can be a machine learning model, such as a Generative Adversarial Network (GAN). The input to the video generation layer can include the corresponding question and generation parameters; the output can include the target operation video. Generation parameters refer to parameters related to the generated target operation video, such as the video speed.
[0153] In some embodiments, the processor may determine the generation parameters based on the user's basic information and usage time. The user may input the user's basic information through the interface. The processor may obtain the user's usage time through the interface.
[0154] User basic information refers to basic information related to the user, such as age, education, and usage experience. Usage experience can include experience with multiple similar software programs. The more experience a user has, the faster the video generation parameters will be.
[0155] In some embodiments, the processor may determine generation parameters by querying a vector database based on basic user information and usage time. The vector database may be constructed based on historical data or may be manually modified and supplemented. The vector database includes reference feature vectors constructed based on historical basic user information and historical usage time, as well as reference generation parameters corresponding to the reference feature vectors.
[0156] In some embodiments, the processor may construct a current feature vector based on the user's basic information and usage time, and by querying a vector database, use the reference generation parameters corresponding to the reference feature vector with the highest similarity to the current feature vector as the generation parameters corresponding to the current user. The similarity may be calculated by calculating Euclidean distance, cosine similarity, or the like.
[0157] The video generation layer may include a first model and a second model. The first model is used to generate a target operation video, which is then input into the second model along with a real teaching video. The second model can then be used to determine whether the data input into the second model is a real teaching video. In some embodiments, the video generation layer may be trained based on a large number of fourth training samples. The fourth training samples may include sample-corresponding questions, sample generation parameters, and real teaching videos. The real teaching video is an operation video that corresponds to the sample-corresponding questions and sample generation parameters and can be used as a reference. In some embodiments, the fourth training samples may be acquired based on historical data.
[0158] The training of the video generation layer consists of multiple stages:
[0159] Phase 1: Fix the parameters of the first model and train the second model. Input the sample-specific question and sample generation parameters into the first model to generate a video. The generated video, along with the sample-specific question and sample generation parameters, is paired together (labeled 0). This pair is then paired with the sample-specific question, sample generation parameters, and the corresponding real-world instructional video to form another data pair (labeled 1). This pair is then used as training data to train the second model, ensuring that it can best distinguish between the generated target operation video and the real-world instructional video.
[0160] The second stage: fix the parameters of the second model and train the first model. The first model and the second model obtained in the first stage are spliced into a composite model. The sample corresponding problem and sample generation parameters are input into the composite model. The composite model outputs a discrimination result (including 0 or 1, 0 indicates that the target operation video output by the first model is not a real teaching video, and 1 indicates that the video output by the first model is a real teaching video). (1-discrimination result) is used as the loss function of the composite model, and the loss function is used to update the parameters of the first model based on the gradient descent method. With the continuous training of the second stage, the more times the composite model outputs the result of 1 or the more times the result of the continuous output is 1, it means that the ability of the first model to output target operation videos similar to real teaching videos is getting stronger and stronger, and the similarity between the target operation video output by the first model and the real teaching video continues to increase.
[0161] Then the first and second stages are cycled, and finally through continuous cycles, the capabilities of the first model and the second model become stronger and stronger, and finally the model converges to obtain a trained video generation layer.
[0162] In some embodiments of this specification, target operation videos are automatically generated based on initial input information and / or target input information, so that the target operation videos can better meet the needs of users, and can quickly and clearly show users the operations that need to be performed. The application scope is wider, the threshold requirements for users are lower, and it helps to improve the user experience.
[0163] Step 230 : Generate operation feedback based on the initial input information and / or the target input information. In some embodiments, step 230 is performed by the processor 140 or the generation module 1530 .
[0164] Operation feedback refers to feedback information related to user input. In some embodiments, operation feedback can include multiple types, such as at least one of image feedback, voice feedback, card feedback, etc.
[0165] For example, the processor can display relevant images to the user in the feedback interface through image feedback.
[0166] For another example, the processor can provide voice feedback to the user. The feedback interface is an interactive interface for displaying feedback operations.
[0167] Figure 7 is an exemplary diagram of voice feedback according to some embodiments of this specification. As shown in Figure 7 , the user's target input information is "Please help me generate an audio message for a clothing promotion announcement." The processor can display the generated audio message as voice feedback on the feedback interface. The feedback interface also includes buttons such as Apply, Try Again, Like, and Text. When the user selects the Text button, the processor can recognize the voice feedback as text.
[0168] Card feedback refers to feedback provided via cards in response to user input. The process of adjusting audio often involves multiple steps, such as creating zones, assigning devices, and playing audio sources. Card feedback allows users to present the complete process or solution in the form of cards within the feedback interface.
[0169] Figure 8 is an exemplary schematic diagram of card feedback according to some embodiments of this specification. As shown in Figure 8, the user's initial input information is "I want to create a partition." The processor can create partition 1 based on the initial input information, detect the devices that can be assigned to partition 1 and the audio sources that can be played, and display the relevant information to the user in the form of a card. As shown in Figure 8, based on "I want to create a partition", the card feedback can feedback the name of the new partition (partition 1), the devices that can be assigned to the new partition (such as CS100_7, CS100_8, CS100_9), and the audio sources that can be used to play in the new partition (such as Spotify). Below the card are an apply button, a try again button, and a like button. If the user selects the apply button for the card, the processor can perform corresponding operations based on the card content. If the user selects the try again button for the card, the processor can re-feed back the card content corresponding to the user input. If the user selects the like button for the card, the processor can mark the card as the user's favorite.
[0170] In some embodiments, the feedback interface for operation feedback may include a user modification window. The user modification window is a window where the user modifies the feedback content. The user can interact with the user modification window by clicking, inputting, etc. to modify the relevant content of the operation feedback.
[0171] Figure 9 is an exemplary diagram of a user modification window according to some embodiments of this specification. As shown in Figure 9, a user can modify the device in the card feedback by clicking the desired device. For example, if the devices in the card feedback include CS100_5, CS100_6, CS100_7, CS100_8, and CS100_9, the user can click CS100_7, CS100_8, or CS100_9 in the user modification window to deselect them, or click CS100_5 or CS100_6 to select the corresponding device.
[0172] Figure 10 is an exemplary diagram of another card feedback interface according to some embodiments of this specification. As shown in Figure 10 , the user's initial input is "Please recommend a party playlist for me." Based on this initial input, the processor can provide a recommended playlist in the form of a card to the feedback interface. The user can directly click on a song to play the corresponding music.
[0173] In some embodiments, operational feedback can also be feedback generated in real time during the execution of the target operation after the user enters the initial input information. For example, operational feedback can be a confirmation message or instruction guidance information that pops up on the interface when performing a display screen operation (such as moving, scaling an operable object, etc.). For another example, operational feedback can be a confirmation message displayed to the user when performing an input error correction operation to confirm whether the error has been corrected. For another example, operational feedback can be a recommended inquiry question displayed in the interface when interacting with intelligent customer service.
[0174] In some embodiments of this specification, by determining operation feedback, relevant content can be fed back to the user in a visual form, making the feedback content clearer and more understandable. Through the user modification window, it helps to further simplify user operations based on user needs and facilitate user use.
[0175] In some embodiments of this specification, the user's target input information is determined based on the user's initial input information and the marking information of the operable object, and the target operation is determined and operation feedback is generated based on the initial input information and / or the target input information, which can effectively simplify the user's operation and improve the user's usage experience and efficiency.
[0176] FIG11 is an exemplary schematic diagram of an intelligent error correction model according to some embodiments of this specification.
[0177] In some embodiments, as described in FIG. 11 , the processor may determine the corrected input information 1150 through the intelligent error correction model 1140 based on the initial input information 1110 and / or the target input information 1120 and perform automatic error correction.
[0178] For more information about the initial input information and the target input information, please refer to the relevant description in FIG2 .
[0179] The intelligent error correction model is a model that corrects initial input information and / or target input information. In some embodiments, the intelligent error correction model can be a machine learning model. For example, the intelligent error correction model can be a neural network model. For another example, the intelligent error correction model can be at least one of a graph neural network (GNN) model, a convolutional neural network (CNN) model, or any combination thereof.
[0180] In some embodiments, the input of the intelligent error correction model may include initial input information and / or target input information; the output may include input information after error correction.
[0181] The error-corrected input information refers to the initial input information and / or target input information after error correction.
[0182] In some embodiments, as shown in FIG11 , the input data of the intelligent error correction model 1140 may be features of the input data represented by an input graph 1130 of a graph structure.
[0183] The input graph can be used to represent various language elements and the relationships between them. In some embodiments, the input graph can be a data structure consisting of nodes and edges, where edges connect nodes and nodes and edges can have features. Language elements refer to elements such as keywords, phrases, and words. As shown in Figure 11, the nodes of the input graph can include A, B, C, etc., and the edges can include AB, BC, etc.
[0184] Nodes correspond to different language elements in the initial input information and / or the target input information. Node features can reflect information related to the language elements. For example, node features can include the word part of speech, domain to which it belongs, and phrase structure characteristics of the language element.
[0185] A word's part of speech refers to the grammatical role or classification a word plays in a sentence. For example, word parts of speech include noun, verb, adjective, adverb, etc.
[0186] Fields reflect the application of a term in a specific field. For example, fields can include daily life, social sciences, humanities, etc.
[0187] Phrase structure features are used to reflect the composition and structural characteristics of a phrase. For example, phrase structure can include subject-predicate phrases, verb-object phrases, parallel phrases, etc.
[0188] Edges can correspond to relationships between language elements. For example, there is an edge between two language elements connected based on grammatical rules in the initial input information and / or target input information. In some embodiments, the edge can be a directed edge, and the grammatical relationship between two adjacent language elements (such as a subject-object relationship, a subject-predicate relationship) is the direction of the edge, with the subject or verb as the starting point of the edge, and the object or predicate as the end point of the edge. Edge features can reflect the dependency relationship, modification relationship, etc. between two adjacent language elements. For example, the features of the edge can include the grammatical relationship between two language elements. Grammatical rules are used to guide the correct construction and use of text, and grammatical rules can combine words and phrases into meaningful sentences. Exemplarily, the "set" and "volume" nodes are connected in sequence to form a directed edge of a subject-object relationship.
[0189] Node and edge features can be determined using various methods based on input data. These methods may be those described in the above embodiments, or other methods may be used. The input data may include current initial input information and / or target input information, or may include historical initial input information and / or target input information.
[0190] In some embodiments, node features may also include the time and place of node input, etc., reflecting the user's usage habits. In some embodiments, node features are also related to the user's input method. For example, when the user inputs the initial input information through text, the corresponding node features may also include the typing speed and the number of backspaces when the user inputs the language element corresponding to the node. When the user inputs the initial input information through voice, the corresponding node features may also include the user's speaking speed corresponding to the node and the user's emotional state. Among them, backspace is used to delete the character before the insertion point. The number of backspaces refers to the total number of times the backspace key is used when the user enters the node.
[0191] In some embodiments, the intelligent error correction model can be obtained based on training data. The training data includes a fifth training sample and a fifth label. For example, the fifth training sample may include a sample input graph, and the fifth label may be whether the word corresponding to the node is wrong, whether the position of the node is wrong, and the corrected node. The nodes and their features, edges and their features of the sample input graph are similar to the above description. The fifth training sample can be determined based on historical data, and the fifth label is determined by a processor or manual annotation. The training process of the intelligent error correction model is similar to that of the prediction model, and reference can be made to the relevant description in FIG12.
[0192] In some embodiments, the processor can connect the corrected nodes and the remaining nodes in sequence based on grammatical rules to form the corrected input information. In some embodiments, the processor can correct the erroneous nodes and connect them in sequence based on the original connection order of the input graph to form the corrected input information. In some embodiments, the initial input information and / or the target input information can be corrected in real time based on the intelligent error correction model.
[0193] In some embodiments of this specification, during the user input process, an intelligent error correction model can efficiently and accurately perform real-time automatic error correction on initial input information and / or target input information, thereby improving the accuracy of the input information, while also enhancing the user experience and improving user input efficiency. Input graphs can be used to correct input information based on grammatical rules, obtaining meaningful and corrected input information, which facilitates the subsequent generation of accurate operational feedback.
[0194] FIG12 is an exemplary schematic diagram of a prediction model according to some embodiments of the present specification.
[0195] In some embodiments, as shown in FIG12 , intelligent customer service interaction includes predicting the user's target input information. The processor may determine a prediction result 1260 through a prediction model 1250 based on the initial input information 1210 , the currently connected device 1220 , the operable object 1230 , and the tag information 1240 ; and determine the target input information 1270 based on the prediction result 1260 .
[0196] For more information about the initial input information, currently connected devices, operable objects, and tag information, please refer to the relevant description of FIG. 2 .
[0197] The prediction result is the predicted content that the user wants to input. The user can choose the prediction result as part or all of the target input information.
[0198] In some embodiments, the number of predicted results may be at least one, and the number of predicted results may be dynamically changing and / or a proportion of the number of times negatively correlated users click on the predicted results.
[0199] Dynamic changes mean that the number of predicted results changes with user input. For example, if a user initially enters "I want to play," the interface displays a preset number of predicted results. If the user doesn't select a result and continues to enter their initial input, this indicates that the number of predicted results is likely too small, and the user can increase the number of predicted results.
[0200] In some embodiments, the number of prediction results needs to meet preset requirements to ensure that the prediction results of the interface can be clearly displayed and convenient for user operation.
[0201] Preset requirements are judgment conditions for evaluating the number of prediction results. For example, the preset requirements may include that the number of prediction results is within a preset range. The preset range can be determined based on experiments or experience. It should be noted that if the number of prediction results in the interactive interface is too large, the interactive interface may become too crowded, making it difficult for users to find the desired prediction results. At the same time, too many prediction results may also cause visual clutter in the interactive interface, reducing the overall aesthetics and user experience. If the number of prediction results in the interactive interface is too small, the interactive interface may appear too simple or empty, lacking necessary functions, making it difficult for users to obtain the desired prediction results, thereby affecting the user experience.
[0202] In some embodiments, the number of predicted results is negatively correlated to the percentage of times users click on the predicted results. For example, the greater the percentage of times users click on the predicted results, the fewer the number of predicted results. The percentage of times users click on the predicted results refers to the ratio of the number of times users click on the predicted results to the number of times they enter the same initial input information within a preset time period. The preset time period can be a system default value or a system preset value. In some embodiments, the processor can determine the percentage of times users click on the predicted results by counting the number of times users click on the predicted results within the preset time period through the network.
[0203] In some embodiments, the number of predicted results may change dynamically, and the number of predicted results is negatively correlated with the percentage of times users click on the predicted results.
[0204] In some embodiments of this specification, by dynamically displaying prediction results to users, prediction results that meet the user's actual needs can be displayed while taking into account the page layout, which is beneficial for users to quickly obtain the content they want to input and further improve the user's usage experience.
[0205] In some embodiments, the prediction results may be displayed in a sorted and / or eye-catching manner.
[0206] In some embodiments, the sorting method refers to sorting based on the evaluation values of multiple prediction results (e.g., ascending or descending order). The evaluation value can be used to evaluate the probability of the prediction result being selected by the user. The higher the evaluation value, the higher the probability that the corresponding prediction result is the content that the user wants to enter.
[0207] In some embodiments, the evaluation value may be determined based on the historical number of clicks on the prediction result. For example, the processor may determine the evaluation value based on the percentage of historical clicks on each prediction result. The percentage of historical clicks on the prediction result refers to the percentage of times users clicked on the prediction result within a historical time period.
[0208] In some embodiments, the eye-catching manner can be implemented in a variety of ways, including but not limited to increasing the font size of the element, using a special font different from the regular font, highlighting the color, underlining, animation effects (e.g., gradually enlarging or rotating the element), etc., or a combination thereof.
[0209] In some embodiments, prediction results with different evaluation values may correspond to one eye-catching method or a combination of multiple eye-catching methods. For example, a prediction result with a higher evaluation value may correspond to a combination of multiple eye-catching methods (e.g., placed at the top of the ranking and with larger font size, etc.), or a prediction result with a higher evaluation value may correspond to one eye-catching method (e.g., only placed at the top of the ranking and with larger font size, or only with larger font size).
[0210] FIG13 is an exemplary schematic diagram of prediction results according to some embodiments of this specification. For example, as shown in FIG13 , after the user enters the initial input information of “i want play”, the prediction results of the interactive interface include 6 words. Assuming that according to the statistical data within the preset time period, the user enters the initial input information of “i want play” a total of 10 times, of which 8 times the user clicked on “aux” (i.e., the user clicked on the prediction result more frequently), the processor will display the prediction result “aux” in a striking manner at the current time, such as by enlarging the font and sorting it first. The current time may refer to the time point when the user enters “i want play”.
[0211] In some embodiments of this specification, the prediction results are displayed in a sorted and / or eye-catching manner, so that some prediction results preferred by the user are relatively forward or relatively eye-catching, allowing the user to quickly obtain the content they want to input, which can further improve the input efficiency.
[0212] In some embodiments, the processor may output a prediction result through a prediction model based on initial input information, currently connected devices, and tag information of operable objects.
[0213] The prediction model can be a natural language processing (NLP) model. The input of the prediction model can include initial input information, currently connected devices, operable objects and their labeling information; the output can be a prediction result.
[0214] In some embodiments, the prediction model can be trained based on a large number of sixth training samples with the sixth label. For example, multiple sixth training samples with the sixth label can be input into the prediction model, a loss function can be constructed using the sixth label and the prediction results of the initial prediction model, and the initial prediction model can be updated based on the iteration of the loss function. Training is completed when the loss function of the initial prediction model meets an end condition, where the end condition can include convergence of the loss function, the number of iterations reaching a threshold, and the like.
[0215] In some embodiments, the processor may first pre-train the prediction model using a large amount of user data, and then perform intensive training on the prediction model using relevant data of the current user.
[0216] In some embodiments, each set of sixth training samples may include initial sample input information, sample operable objects and their labeling information, and sample connection devices. The sixth training samples may be acquired based on historical data. The sixth label is the historical actual transmission information corresponding to the sixth training sample. The sixth label may be manually annotated.
[0217] In some embodiments, the sixth training sample and the sixth label can be obtained through annotation. For example, a large number of users' complete input sentences can be used as the sixth label; keywords in the complete input sentences can be extracted as sample initial input information; and the operable objects and their labeling information and connected devices corresponding to each complete input sentence can be determined as sample operable objects and their labeling information and sample connected devices.
[0218] In some embodiments, when performing enhanced training on the prediction model, the sixth training sample may further include personalized information about the current user. For example, basic information such as the user's face, identity, and age may be included. By incorporating personalized user information into the sample, the model's personalized prediction capabilities can be enhanced.
[0219] In some embodiments, the processor can determine the target input information based on the prediction results in various ways. For example, the processor can capture the user's selection operation on the prediction result in the interactive interface through the network. Specifically, the processor can obtain the user's click operation and / or gesture operation on the prediction result in the interactive interface, and determine the prediction result selected by the user as part or all of the information in the target input information. In some embodiments, based on the initial input information entered by the user in the interactive interface, combined with the operable objects and tag information of the interactive interface, the application software (or background processor) can automatically predict the user's prediction result and display it in the input information display box of the interactive interface. Furthermore, based on the prediction result, further content can be displayed to the user in the interactive interface. As shown in Figure 13, the user's initial input information in the interactive interface is "I want to play". The processor can obtain the playback content such as "aux, USB, playlist, Roon, AirPlay, Spotify" based on the prediction result of the initial input information and display it to the user. The processor can use the prediction result selected by the user through the interface as part or all of the information in the target input information.
[0220] In some embodiments of this specification, the prediction model can accurately and efficiently predict the content that the user wants to input, which is conducive to improving the efficiency of the user's input and further improving the user's usage experience.
[0221] FIG14 is an exemplary schematic diagram of a prediction recommendation model according to some embodiments of this specification.
[0222] In some embodiments, intelligent customer service interaction includes predicting the user's emotional state. The processor can predict the emotional state based on the user's input characteristics and make function recommendations based on the emotional state.
[0223] The user's emotional state refers to a state that can reflect the user's current mood. For example, the categories of the user's emotional state may include happy, sad, angry, depressed, stable, etc.
[0224] Function recommendation refers to the process of recommending relevant functions or services to users based on their needs and status. In some embodiments, different emotional states may correspond to different function recommendations.
[0225] Input features are used to reflect the characteristics of the user's initial input information. In some embodiments, the user's initial input information is voice. Correspondingly, the input features may include indicators such as pitch, speaking rate, frequency, and energy. Energy refers to the strength of the sound, that is, the loudness of the sound. Energy can be determined by the amplitude of the sound. The larger the amplitude, the greater the energy and the louder the sound. In some embodiments, the processor can determine the energy of the sound based on the sum of the square powers of the sound amplitude within a time period. The time period can be determined based on the duration of the sound.
[0226] In some embodiments, the input features can be represented by a feature vector as (A, B, C, ...), where element A represents pitch, B represents speaking rate, C represents energy, and so on.
[0227] In some embodiments, the processor can obtain the user's audio data through an installed recording device and then extract the user's input features from the audio data. The audio data can refer to the voice of the user interacting with the intelligent customer service.
[0228] In some embodiments, the processor may extract input features using an audio feature extraction algorithm. The audio feature extraction algorithm may include, but is not limited to, linear prediction coefficients (LPC), perceptual linear prediction coefficients (PLP), linear prediction cepstral coefficients (LPCC), and mel-frequency cepstrum coefficients (MFCC).
[0229] In some embodiments, before extracting the input features, the processor may pre-process the audio data. The pre-processing of the audio data may include at least one of pre-emphasis, framing, and windowing.
[0230] In some embodiments, the processor may predict the user's emotional state based on the user's input features, and obtain a recommended function (such as making a phone call, listening to a song, sending a text message, etc.) corresponding to the current user's emotional state based on the corresponding relationship between the emotional states of different users and different recommended functions. The corresponding relationship can be determined through historical data or prior knowledge.
[0231] In some embodiments, as shown in FIG. 14 , the processor may process the user's input features 1410 based on a prediction recommendation model 1420 to predict the user's emotional state 1430 and a list of recommended functions 1440 corresponding to the emotional state.
[0232] The recommended function list is a list of different recommended functions. Recommended functions are functions recommended for users to use. For example, if a user displays positive emotions, this indicates that they prefer functions that provide entertainment or relaxation. The processor may identify functions such as listening to music or playing short videos as recommended functions. If a user displays negative emotions, this indicates that they prefer functions that help them solve problems or relieve stress. The processor may identify functions such as making phone calls as recommended functions.
[0233] In some embodiments, the prediction and recommendation model may be a machine learning model. For example, the prediction and recommendation model may include any one or a combination of various feasible models, such as a recurrent neural network (RNN) model, a deep neural network (DNN) model, or a convolutional neural network (CNN) model.
[0234] In some embodiments, the input of the predictive recommendation model may include input features (the input features may be represented by feature vectors), and the output may include the user's emotional state and a list of recommended functions applicable to the emotional state. The output user's emotional state may be represented by an emotional state category.
[0235] In some embodiments, the predictive recommendation model can be trained using various feasible methods based on a large number of seventh training samples with the seventh label. For example, parameter updates can be performed using a gradient descent method. The training process of the predictive recommendation model is similar to that of the prediction model, and can be seen in the relevant description of FIG12 .
[0236] In some embodiments, the seventh training sample may include a sample input feature of a sample user. The seventh training sample may be obtained based on historical data.
[0237] In some embodiments, the seventh label can be the category of the sample user's actual emotional state or the actual function of the sample user's historical operation. For example, the fifth label is obtained by labeling the function selected by the sample user during actual operation and the category of the sample user's actual emotional state.
[0238] In some embodiments, the categories of actual emotional states can be diverse. For example, the categories of actual emotional states can be happy, sad, angry, depressed, stable, etc. For another example, the categories of actual emotional states can be divided into certain levels. For example, the emotional classification of happiness can be divided into first-level happiness, second-level happiness, third-level happiness, etc.
[0239] In some embodiments, the processor may obtain image or video data of the sample user's expression through an installed camera, and manually label the category of the sample user's actual emotional state based on the image and / or video data. The camera may include a camera, a still camera, etc.
[0240] In some embodiments of this specification, by predicting emotional states and recommending different functions based on user input features, the user's internal emotions and external expressions are fully considered, enabling targeted recommendations of different functions for different users, thereby improving the user experience during input. By using models to process input features, the type and number of functions to be pushed can be more accurately determined.
[0241] FIG15 is an exemplary module diagram of a natural language-based interaction system according to some embodiments of the present specification.
[0242] In some embodiments, the natural language-based interaction system 1500 may include a first determination module 1510 , a second determination module 1520 , and a generation module 1530 .
[0243] In some embodiments, the first determination module 1510 may be configured to determine the user's target input information based on the user's initial input information and tag information of the operable object, wherein the tag information reflects the characteristics of the operable object.
[0244] In some embodiments, the second determination module 1520 can be configured to determine a target operation based on the initial input information and / or the target input information, where the target operation includes at least one of a display screen operation, an input error correction operation, and an intelligent customer service interaction.
[0245] In some embodiments, the generation module 1530 may be configured to generate operation feedback based on the initial input information and / or the target input information, and the feedback interface of the operation feedback includes a user modification window.
[0246] In some embodiments, the natural language-based interaction system 1500 further includes a drag module (not shown in the figure), which may be configured to determine target input information based on a user's drag instruction on an operation target in the display screen.
[0247] In some embodiments, the operation target includes at least one of a draggable target and a non-draggable target, and the first display modes corresponding to the draggable target and the non-draggable target are different.
[0248] In some embodiments, the display screen comes from at least one platform, and the at least one platform corresponds to the second display mode of the operation target; and / or the display screen comes from at least one system, and the at least one system corresponds to the third display mode of the operation target, and the second display mode and the third display mode are different.
[0249] In some embodiments, the second determination module 1520 is further configured to: based on the initial input information and / or target input information, determine the corrected input information through an intelligent error correction model, and perform automatic error correction, where the intelligent error correction model is a machine learning model.
[0250] In some embodiments, intelligent customer service interaction includes predicting the user's target input information, and the second determination module 1520 is further configured to: determine the prediction result through a prediction model based on the initial input information, the currently connected device, the operable object, and the tag information, where the prediction model is a machine learning model; and determine the target input information based on the prediction result.
[0251] In some embodiments, the number of predicted results is at least one, and the number of predicted results changes dynamically and / or the proportion of times negatively correlated users click on the predicted results.
[0252] In some embodiments, the intelligent customer service interaction includes predicting the user's estimated usage function, and the second determination module 1520 is further configured to: determine the user's user characteristics based on the user's behavioral habits; and predict the estimated usage function based on the user characteristics.
[0253] In some embodiments, the intelligent customer service interaction includes predicting the user's emotional state, and the second determination module 1520 is further configured to: predict the emotional state based on the user's input characteristics, and recommend functions based on the emotional state.
[0254] For more information about the natural language-based interactive system 1500 , please refer to the above description.
[0255] It should be understood that the natural language-based interactive system 1500 and its modules shown in Figure 15 can be implemented in various ways. It should be noted that the above description of the natural language-based interactive system 1500 and its modules is for convenience only and does not limit this specification to the scope of the embodiments illustrated. It is understood that for those skilled in the art, after understanding the principles of the system, it is possible to arbitrarily combine the various modules or form subsystems connected to other modules without departing from these principles. In some embodiments, the first determination module 1510, the second determination module 1520, and the generation module 1530 disclosed in Figure 15 can be different modules in a system, or a single module can implement the functions of two or more of the above-mentioned modules. For example, the modules can share a storage module, or each module can have its own storage module. Variations such as these are within the scope of protection of this specification.
[0256] Some embodiments of the present specification provide a natural language-based interaction device, comprising at least one processor and at least one memory, wherein the at least one memory is used to store computer instructions, and the at least one processor is used to execute at least part of the computer instructions to implement the aforementioned natural language-based interaction method.
[0257] Some embodiments of this specification provide a computer-readable storage medium, which stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the aforementioned natural language-based interaction method.
[0258] It should be noted that the above description of the relevant processes is for illustration and purpose only and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the processes under the guidance of this specification. However, such modifications and changes are still within the scope of this specification.
[0259] While the basic concepts have been described above, it will be apparent to those skilled in the art that the detailed disclosure is merely illustrative and does not limit this specification. Although not explicitly stated herein, various modifications, improvements, and revisions to this specification may be made by those skilled in the art. Such modifications, improvements, and revisions are suggested in this specification and remain within the spirit and scope of the exemplary embodiments of this specification.
[0260] This specification also uses specific terms to describe the embodiments of this specification. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "one embodiment," "an embodiment," or "an alternative embodiment" two or more times in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics of one or more embodiments of this specification may be appropriately combined.
[0261] In addition, unless expressly stated in the claims, the order of the processing elements and sequences described in this specification, the use of alphanumeric characters, or the use of other names are not intended to limit the order of the processes and laminar flow hoods in this specification. Although the above disclosure discusses some of the invention embodiments currently considered useful through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the spirit and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.
[0262] Similarly, it should be noted that, in order to simplify the presentation of this specification and thus facilitate understanding of one or more embodiments of the invention, the foregoing descriptions of the embodiments of this specification sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not imply that the subject matter of this specification requires more features than those recited in the claims. In fact, an embodiment may have fewer features than all of the features of a single disclosed embodiment.
[0263] In some embodiments, numbers are used to describe the quantity of components and attributes. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may change according to the required characteristics of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the settings of such numerical values are as accurate as possible within the feasible range.
[0264] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, and documents, cited in this specification is hereby incorporated by reference in its entirety. This includes application history documents that are inconsistent with or conflict with the content of this specification, as well as documents (currently or subsequently attached to this specification) that limit the broadest scope of the claims of this specification. It should be noted that if the descriptions, definitions, and / or terminology used in the accompanying materials are inconsistent or conflicting with the content of this specification, the descriptions, definitions, and / or terminology used in this specification will control.
[0265] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.
Claims
1. A natural language based interaction method, characterized in that: The method is executed by a processor, comprising: Determining target input information of the user based on initial input information of the user and tag information of the operable object, wherein the tag information reflects the characteristics of the operable object; Determine a target operation based on the initial input information and / or the target input information, where the target operation includes at least one of a screen display operation, an input error correction operation, and an intelligent customer service interaction; and Based on the initial input information and / or the target input information, an operation feedback is generated, and a feedback interface of the operation feedback includes a user modification window.
2. The method according to claim 1, characterized in that The method further comprises: The target input information is determined based on a drag instruction of the user to an operation target in the display screen.
3. The method according to claim 2, characterized in that The operation target includes at least one of a draggable target and a non-draggable target, and the first display modes corresponding to the draggable target and the non-draggable target are different.
4. The method according to claim 2, characterized in that: The display screen originates from at least one platform, and the at least one platform corresponds to the second display mode of the operation target; and / or The display screen originates from at least one system, the at least one system corresponds to a third display mode of the operation target, and the second display mode is different from the third display mode.
5. The method according to claim 1, characterized in that The input error correction operation includes: Based on the initial input information and / or the target input information, the input information after error correction is determined by an intelligent error correction model, and automatic error correction is performed. The intelligent error correction model is a machine learning model.
6. The method according to claim 1, characterized in that The intelligent customer service interaction includes predicting the target input information of the user, and the method further includes: Determine a prediction result by a prediction model based on the initial input information, the currently connected device, the operable object, and the tag information, wherein the prediction model is a machine learning model; Based on the prediction result, the target input information is determined.
7. The method according to claim 6, characterized in that The number of the predicted results is at least one, and the number of the predicted results changes dynamically and / or is negatively correlated with the proportion of times the user clicks on the predicted results.
8. The method according to claim 6, characterized in that The prediction results are displayed in a sorted manner and / or a striking manner.
9. The method according to claim 1, characterized in that: The intelligent customer service interaction includes predicting the estimated usage function of the user, and the method further includes: Determining user characteristics of the user based on user behavior habits; The estimated usage function is predicted based on the user characteristics.
10. The method according to claim 1, characterized in that The intelligent customer service interaction includes predicting the emotional state of the user, and the method further includes: The emotional state is predicted based on the input features of the user, and function recommendations are made based on the emotional state.
11. An interactive system based on natural language, characterized in that: The system comprises: A first determination module is configured to determine target input information of the user based on initial input information of the user and tag information of the operable object, wherein the tag information reflects the characteristics of the operable object; A second determination module is configured to determine a target operation based on the initial input information and / or the target input information, wherein the target operation includes at least one of a screen display operation, an input error correction operation, and an intelligent customer service interaction; and The generating module is configured to generate operation feedback based on the initial input information and / or the target input information, wherein the feedback interface of the operation feedback includes a user modification window.
12. The system according to claim 11, characterized in that The system further includes a drag module, wherein the drag module is configured to: The target input information is determined based on a drag instruction of the user to an operation target in the display screen.
13. The system according to claim 12, characterized in that The operation target includes at least one of a draggable target and a non-draggable target, and the first display modes corresponding to the draggable target and the non-draggable target are different.
14. The system according to claim 12, characterized in that The display screen originates from at least one platform, and the at least one platform corresponds to the second display mode of the operation target; and / or The display screen originates from at least one system, the at least one system corresponds to a third display mode of the operation target, and the second display mode is different from the third display mode.
15. The system according to claim 11, characterized in that The second determining module is further configured to: Based on the initial input information and / or the target input information, the input information after error correction is determined by an intelligent error correction model, and automatic error correction is performed. The intelligent error correction model is a machine learning model.
16. The system according to claim 11, characterized in that The intelligent customer service interaction includes predicting the target input information of the user, and the second determination module is further configured to: Determine a prediction result by a prediction model based on the initial input information, the currently connected device, the operable object, and the tag information, wherein the prediction model is a machine learning model; Based on the prediction result, the target input information is determined.
17. The system according to claim 16, characterized in that The number of the predicted results is at least one, and the number of the predicted results changes dynamically and / or is negatively correlated with the proportion of times the user clicks on the predicted results.
18. The system according to claim 11, characterized in that The intelligent customer service interaction includes predicting the estimated usage function of the user, and the second determination module is further configured to: Determining user characteristics of the user based on user behavior habits; The estimated usage function is predicted based on the user characteristics.
19. The system according to claim 11, characterized in that The intelligent customer service interaction includes predicting the emotional state of the user, and the second determination module is further configured to: The emotional state is predicted based on the input features of the user, and function recommendations are made based on the emotional state.
20. A computer-readable storage medium, characterized in that: The storage medium stores computer instructions. When the computer reads the computer instructions in the storage medium, the computer executes the natural language-based interaction method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Voice input error correction method and device based on artificial intelligence
CN107678561A
Playing processing method and device, equipment and storage medium
CN110297940A
Intelligent information exchange method, system and device and readable storage medium
CN115408608A
Information interaction method and device, computer equipment and storage medium
CN117076627A
Human-computer interaction method, and electronic device and storage medium thereof
US20210065682A1