Inference method and computing device
By acquiring and analyzing users' paralinguistic information, especially stressed words and emotional features, the inference results of the large model were adjusted, which solved the shortcomings of the large model in understanding user input and improved the user experience.
Patent Information
- Application Number
- CN202510884874.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-11
AI Technical Summary
Large models often struggle to fully and accurately understand user input, resulting in outputs that fail to meet user expectations and diminish the user experience.
By acquiring paralinguistic information from reasoning tasks, including acoustic and perceptual features, we can determine the user's emotions and concerns. We can then adjust the reasoning results using stress word reasoning models and emotion reasoning models, and optimize the output by combining knowledge graphs and response templates.
It improves the accuracy of the computing device's output, enhancing user satisfaction and experience.
Smart Images

Figure CN120930776A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a reasoning method and computing device. Background Technology
[0002] With the continuous development of large-scale model technology, intelligent services based on large-scale models, such as speech synthesis and intelligent question answering, are receiving increasing attention. These intelligent services provide users with a convenient and efficient experience through large-scale model technology.
[0003] Currently, when users use large models for intelligent question answering, the large models have difficulty fully and accurately understanding the questions entered by users, resulting in the output of the large models failing to meet the user's expectations, thereby reducing the user's experience. Summary of the Invention
[0004] This application provides a reasoning method and a computing device, which can effectively improve user satisfaction with the reasoning results output by the computing device and enhance the user experience.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0006] In a first aspect, a reasoning method is provided, applied to a computing device. The method includes: responding to a reasoning task input by a user, acquiring paralinguistic information contained in the reasoning task, processing the reasoning task to obtain a first reasoning result, adjusting the first reasoning result based on the paralinguistic information to obtain a second reasoning result, and outputting the second reasoning result.
[0007] The second inference result matches the paralinguistic information more closely than the first inference result. Paralinguistic information is used to represent the user's emotions and / or focus.
[0008] Through the above technical solution, when a computing device performs reasoning on a user-input reasoning task, it can take into account information such as the user's emotions and / or concerns implicit in the reasoning task, and thus adjust the reasoning result based on the user's emotions and / or concerns. In this way, the reasoning result output by the computing device (such as the second reasoning result) can better match the user's expectations, effectively improving user satisfaction with the reasoning result output by the computing device and enhancing the user experience.
[0009] In one alternative implementation, the reasoning task may include audio information. Accordingly, obtaining the paralinguistic information contained in the reasoning task may specifically include: determining the paralinguistic information based on acoustic features in the audio information.
[0010] Through the above technical solution, the computing device can determine the sub-language information contained in the audio information based on the acoustic features in the audio information. This not only enhances the feasibility of the solution, but also effectively improves the efficiency of the computing device in determining the sub-language information.
[0011] In one optional implementation, the acoustic features mentioned above may include physical features, which refer to the objective attributes of audio information, and the paralinguistic information includes stressed words in the audio information. Accordingly, the above-mentioned determination of paralinguistic information based on acoustic features in audio information may specifically include: determining stressed words based on physical features.
[0012] Through the above technical solution, the computing device can identify the stressed words in the audio information. Stressed words can usually represent the user's focus. In this way, the computing device can obtain the user's focus through stressed words, so that the computing device can adjust the inference results based on the user's focus to meet the user's needs, improve the user's satisfaction with the inference results, and enhance the user's experience.
[0013] In one alternative implementation, a stress word inference model may be deployed in the computing device. This model is used to determine stress based on physical features. Accordingly, determining stress words based on physical features may specifically include: inputting audio information into the stress word inference model to obtain the stress words.
[0014] Through the above technical solution, the computing device can quickly obtain the stressed words in the audio information through the stressed word inference model. In this way, the efficiency of the computing device in determining stressed words can be effectively improved, thereby effectively improving the efficiency of the computing device in reasoning for the inference task input by the user.
[0015] In one optional implementation, the acoustic features may include perceptual features, which characterize the user's subjective feelings about the audio information, and the paralinguistic information includes emotional features in the audio information. Accordingly, the above-mentioned determination of paralinguistic information based on acoustic features may specifically include: determining emotional features based on perceptual features.
[0016] Through the above technical solution, the computing device can determine the emotional features in the audio information. These emotional features can usually represent the user's emotions. In this way, the computing device can obtain the user's emotions, so that the subsequent computing device can adjust the inference results based on the user's emotions to meet the user's needs, improve the user's satisfaction with the inference results, and enhance the user's experience.
[0017] In one optional implementation, the above-mentioned adjustment of the first inference result based on the paralinguistic information to obtain the second inference result may specifically include: determining a response template that matches the paralinguistic information, and using the response template to adjust the first inference result to obtain the second inference result.
[0018] Through the above technical solution, the computing device can adjust the reasoning result by using a response template that matches the paralinguistic information contained in the reasoning task, so that the adjusted reasoning result (such as the second reasoning result) can better meet the user's needs, thereby improving the user's satisfaction with the reasoning result and enhancing the user's experience.
[0019] In one optional implementation, the computing device may store multiple reference tasks and reference results corresponding to each reference task. Before processing the inference task to obtain the first inference result, the computing device may further determine a first entity contained in the inference task. Accordingly, processing the inference task to obtain the first inference result may specifically include: determining the target reference task with the highest similarity to the inference task from multiple reference tasks, obtaining entity information of the second entity, and determining the first inference result based on the reference result corresponding to the target reference task and the entity information of the second entity.
[0020] The second entity refers to an entity that has a specific relationship with the first entity.
[0021] As can be seen from the above technical solution, the first reasoning result not only includes the result obtained after reasoning about the reasoning task, but also includes entity information of the second entity associated with the first entity in the reasoning task. In this way, the content of the first reasoning result can be richer and more comprehensive.
[0022] In one optional implementation, a knowledge graph (KG) may be deployed in the computing device. The knowledge graph stores entities that are associated with different entities, as well as entity information for each entity. Accordingly, obtaining the entity information of the second entity may specifically include: determining the second entity from the knowledge graph and obtaining the entity information of the second entity.
[0023] The above technical solution can effectively improve the efficiency of computing devices in obtaining entity information of a second entity.
[0024] In one alternative implementation, before determining the first entity included in the reasoning task, the computing device may further modify the reasoning task. Accordingly, determining the first entity included in the reasoning task may specifically include: determining the first entity from the modified reasoning task.
[0025] The above technical solution can correct the reasoning task, which can not only effectively improve the accuracy of extracting the first entity, but also effectively improve the accuracy of subsequent computing devices in determining the target reference task.
[0026] In one optional implementation, the reasoning task may include text information. Accordingly, obtaining the paralinguistic information contained in the reasoning task may specifically include: determining keywords in the text information, and determining paralinguistic information based on the keywords.
[0027] Keywords may include at least one of the following: words used to represent entities, words used to represent user emotions, or words that appear more frequently than a threshold.
[0028] Through the above technical solution, the computing device can determine the paralinguistic information contained in the audio information based on the keywords in the text information. This not only enhances the feasibility of the solution, but also effectively improves the efficiency of the computing device in determining the paralinguistic information.
[0029] Secondly, an inference apparatus is provided, comprising: functional units for performing any of the methods provided in the first aspect, wherein the actions performed by each functional unit are implemented by hardware or by hardware executing corresponding software. For example, a model training apparatus may include: an acquisition unit, a processing unit, an adjustment unit, and an output unit.
[0030] The system includes an acquisition unit for acquiring paralinguistic information contained in the inference task in response to user input. A processing unit processes the inference task to obtain a first inference result. An adjustment unit adjusts the first inference result based on the paralinguistic information to obtain a second inference result. An output unit outputs the second inference result.
[0031] Thirdly, a computing device is provided, comprising: a processor and a memory, the processor being connected to the memory. The memory is used to store computer-executable instructions, and the processor executes the computer-executable instructions stored in the memory, thereby implementing any of the methods provided in the first aspect.
[0032] Fourthly, a chip is provided, comprising: a processor and an interface circuit; the interface circuit is used to receive code instructions and transmit them to the processor; the processor is used to execute the code instructions to perform any of the methods provided in the first aspect above.
[0033] Fifthly, a computer-readable storage medium is provided, storing computer-executable instructions that, when executed on a computer, cause the computer to perform any of the methods provided in the first aspect above.
[0034] In a sixth aspect, a computer program product is provided, including computer execution instructions that, when executed on a computer, cause the computer to perform any of the methods provided in the first aspect above.
[0035] The technical effects of any of the implementation methods in aspects two through six can be found in the technical effects of different implementation methods in aspect one, and will not be repeated here. Attached Figure Description
[0036] Figure 1 This application provides a schematic diagram of the architecture of a communication system.
[0037] Figure 2 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0038] Figure 3 A flowchart illustrating a reasoning method provided in an embodiment of this application;
[0039] Figure 4 A schematic flowchart illustrating the process of determining a first reasoning result, provided as an embodiment of this application;
[0040] Figure 5 A schematic diagram of a software module included in a computing device according to an embodiment of this application;
[0041] Figure 6 A flowchart illustrating another reasoning method provided in an embodiment of this application;
[0042] Figure 7 This is a schematic diagram of the structure of an inference device provided in an embodiment of this application. Detailed Implementation
[0043] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0044] In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is merely a description of the relationship between the related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. A and B can be singular or plural.
[0045] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0046] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" are not necessarily different. Meanwhile, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is being used as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.
[0047] The following describes the terminology used in the embodiments of this application.
[0048] Large language model (LLM): LLM mainly refers to deep learning models, such as bidirectional encoder representations from transformers (BERT) and generative pretrained transformers (GPT). LLM can handle complex tasks such as language understanding and text generation.
[0049] Knowledge base: A knowledge base is a database that stores large amounts of data that computers can use to search, retrieve, and answer user questions.
[0050] Multimodal: Multimodal refers to different types of data (such as text, images, audio, video, etc.). In the field of artificial intelligence, multimodal learning refers to the ability of machine learning algorithms to learn from multiple types of data and fuse different types of data to improve model performance.
[0051] Automatic speech recognition (ASR) is a technology that converts human speech into text. ASR can be widely used in various scenarios, such as voice assistants, telephone customer service systems, and voice input methods.
[0052] Vector database: A vector database is a database used to store and query vector data. It can be used in scenarios involving similarity search, such as image recognition, recommendation systems, or natural language processing.
[0053] Natural Language Processing (NLP): NLP aims to enable computers to understand, interpret, and generate human natural language. The goal of NLP is to build intelligent systems that can communicate effectively with humans, not only parsing text but also responding based on context.
[0054] Mel-frequency cepstral coefficients (MFCC): MFCC is a feature extraction method widely used in audio processing.
[0055] A tokenization model can divide a continuous natural language text into individual lexical units (tokens), or words.
[0056] The application scenarios of the embodiments of this application will be described below by way of example.
[0057] With the continuous development of large model technology, intelligent services based on large models, such as speech synthesis and intelligent question answering, are being used more and more widely. These intelligent services provide users with a convenient and efficient experience through large model technology.
[0058] This application provides a reasoning method applied to a computing device. The computing device can respond to a reasoning task input by a user, acquire paralinguistic information contained in the reasoning task to characterize the user's emotions and / or points of interest, process the reasoning task, and obtain a first reasoning result. Subsequently, the computing device can adjust the first reasoning result based on the paralinguistic information to obtain a second reasoning result, and output the second reasoning result.
[0059] Through the above technical solution, when a computing device performs reasoning on a user-input reasoning task, it can take into account information such as the user's emotions and / or concerns implicit in the reasoning task, and thus adjust the reasoning result based on the user's emotions and / or concerns. In this way, the reasoning result output by the computing device (such as the second reasoning result) can better match the user's expectations, effectively improving user satisfaction with the reasoning result output by the computing device and enhancing the user experience.
[0060] The system architecture of the embodiments of this application will be described below by way of example.
[0061] Figure 1 This is a schematic diagram of the architecture of a communication system provided in an embodiment of this application. Figure 1 As shown, the communication system may include a terminal device 101 and a computing device 102. The terminal device 101 and the computing device 102 are communicatively connected.
[0062] Terminal equipment 101 may also be referred to as user equipment (UE) or terminal equipment (TE). For example, terminal equipment may include personal digital assistant (PDA), ultra-mobile personal computer (UMPC), laptop, netbook, desktop computer, all-in-one computer, mobile phone, tablet, in-vehicle equipment, or wearable device, etc.
[0063] The computing device 102 can be a network device. A network device can include servers, etc. A server can be a single physical server, or two or more physical servers sharing different responsibilities and working together to achieve various server functions, or a virtual server (also called a virtual machine) running on a physical server, etc. For example, a server can be a blade server, a high-density server, a rack server, or a tower server, etc.
[0064] In this embodiment, when a user intends to use artificial intelligence (AI) for intelligent question-answering services, they can input a corresponding reasoning task (also called a question) into their terminal device 101. After receiving the reasoning task input by the user, the terminal device 101 can send the reasoning task to the computing device 102. After receiving the reasoning task, the computing device can obtain the paralinguistic information contained in the reasoning task and process the reasoning task to obtain a reasoning result (i.e., a first reasoning result). Subsequently, the computing device 102 can adjust the first reasoning result based on the paralinguistic information contained in the reasoning task to obtain a second reasoning result.
[0065] The computing device 102 can send the second inference result to the terminal device 101. After receiving the second inference result, the terminal device 101 can output the second inference result, for example, by displaying the second inference result on its own screen or by playing the second inference result through a speaker. Here, there is no specific limitation on how the terminal device 101 outputs the second inference result.
[0066] The second reasoning result matches the paralinguistic information contained in the reasoning task better than the first reasoning result matches the paralinguistic information contained in the reasoning task.
[0067] It should be noted that the embodiments of this application do not limit the device form of the computing device 102. The following uses a server as an example to describe the system architecture of the computing device 102 provided in the embodiments of this application.
[0068] Figure 2 This is a schematic diagram of a computing device 102 provided in an embodiment of this application. Figure 2 As shown, the computing device 102 includes a processor 202 and memory 204. The processor 202 is connected to the memory 204 via a double data rate (DDR) bus 203. Here, the DDR bus 203 can also be replaced with other types of buses; this embodiment does not limit the bus type. In addition, the computing device 102 also includes various input / output (I / O) devices 207, which the processor 202 can access via a peripheral component interconnect express (PCIe) bus 205.
[0069] Processor 202 is the computing and control core of computing device 102. Processor 202 may include one or more processor cores 201. Processor 202 may be a very large-scale integrated circuit. An operating system and other software programs are installed in processor 202, enabling processor 202 to access memory 204 and various PCIe devices.
[0070] It is understood that, in this embodiment of the invention, the core 201 in the processor 202 may be, for example, a central processing unit (CPU) or another application-specific integrated circuit (ASIC). The processor 202 may also be another general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In practical applications, the computing device 102 may also include multiple processors.
[0071] The memory controller is a bus circuit controller within the computing device 102 that controls the memory 204 and manages and plans the data transfer between the memory 204 and the core 201. Data can be exchanged between the memory 204 and the core 201 through the memory controller. The memory controller can be a separate chip and is connected to the core 201 via the system bus.
[0072] Those skilled in the art will understand that the memory controller can be integrated into the processor 202, built into the northbridge, or be a separate memory controller chip. This embodiment of the invention does not limit the specific location or form of the memory controller. In practical applications, the memory controller can control the necessary logic to write data to or read data from memory 204. The memory controller 204 can be a memory controller in a processor system such as a general-purpose processor, a dedicated accelerator, a GPU, an FPGA, or an embedded processor.
[0073] Memory 204 is the main memory of computing device 102. Memory 204 is typically used to store various running software in the operating system, input and output data, and information exchanged with external storage. To improve the access speed of processor 202, memory 204 needs to have the advantage of high access speed. In traditional computer system architectures, dynamic random access memory (DRAM) is usually used as memory 204. Processor 202 can access memory 204 at high speed through memory controller, performing read and write operations on any storage unit in memory 204. In addition to DRAM, memory 204 can also be other random access memory, such as static random access memory (SRAM). Alternatively, memory 204 can also be read-only memory (ROM). For example, read-only memory can be programmable read-only memory (PROM) or erasable programmable read-only memory (EPROM). This embodiment does not limit the number or type of memory 204. Furthermore, memory 204 can be configured to have a power-saving function. The power-saving function means that the data stored in the memory will not be lost when the system experiences a power outage and is then powered on again. Memory 204 with a power-saving function is called non-volatile memory.
[0074] I / O device 207 refers to hardware capable of data transmission, or a device that interfaces with an I / O interface. Common I / O devices include network cards, printers, keyboards, and mice. All external storage devices can also be used as I / O devices, such as hard drives, floppy disks, and optical discs. The processor 202 can access each I / O device 207 via the PCIe bus 205. It should be noted that the PCIe bus 205 is just one example and can be replaced with other buses, such as the unified bus (UB) bus.
[0075] The baseboard management controller (BMC) 206 can perform firmware upgrades, manage the device's operating status, and troubleshoot faults when the computing device 102 is not powered on. The processor 202 can access the baseboard management controller 206 via the PCIe bus 205. The baseboard management controller 206 can also be connected to at least one sensor. The sensor acquires status data of the computing device 102, including temperature data, current data, voltage data, etc. No specific limitations are made on the type of status data in this application. The baseboard management controller 206 communicates with the processor 202 via the PCIe bus or other types of buses, for example, by transmitting the acquired status data to the processor 202 for processing. The baseboard management controller 206 can also maintain the program code in memory, including upgrading or restoring it. The baseboard management controller 206 can also control the power supply circuit or clock circuit within the computing device 102. In summary, the baseboard management controller 206 can manage the computing device 102 in the above ways. However, the baseboard management controller 206 is only an optional device. In some implementations, the processor 202 can communicate directly with the sensors, thereby directly managing and maintaining the computing device 102.
[0076] It should be noted that the system architecture and application scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0077] The following are Figure 1 and Figure 2 Taking the computing device shown in the figure as an example, the inference method provided in the embodiments of this application will be described in detail. Figure 3 A flowchart illustrating a reasoning method provided in an embodiment of this application is shown below. Figure 3 As shown, the method includes the following steps S301-S304.
[0078] S301, in response to the user's input reasoning task, obtains the paralinguistic information contained in the reasoning task.
[0079] Paralinguistic information is used to represent the user's emotions and / or focus. These emotions can include, but are not limited to, different feelings such as happiness, anger, resentment, anxiety, or joy.
[0080] This application does not specifically limit the type of inference task in its embodiments. In one example, the inference task can be text information input by the user, such as tomorrow's weather forecast, or text information on how to use a rice cooker. In another example, the inference task can be image information input by the user. In yet another example, the inference task can be audio or video information input by the user.
[0081] Specifically, taking a mobile phone as an example, the user's terminal device can be equipped with an artificial intelligence (AI) question-answering platform, which provides intelligent question-answering services. Based on this, when a user intends to use AI for intelligent question-answering services, they can launch the AI question-answering platform on their phone and log in to their account. After logging in, the user can input corresponding inference tasks on the AI question-answering platform according to their needs. Once the mobile phone receives the inference task, it can send the inference task to the computing device. After receiving the inference task, the computing device can obtain the paralinguistic information contained in the inference task.
[0082] In this embodiment, the AI question-answering platform is not specifically limited. In one example, the AI question-answering platform can be AI software installed on a terminal device that can provide intelligent question-answering services. In this case, the process of the user launching the AI question-answering platform on their mobile phone is as follows: an icon for the AI software can be displayed on the phone's desktop, and the user can click the icon to launch the AI software. In another example, the AI question-answering platform can be an AI website that provides intelligent question-answering services. In this case, the process of the user launching the AI question-answering platform on their mobile phone is as follows: the user can enter the URL of the AI website in their mobile browser to log in to the AI website.
[0083] For example, using paralinguistic information to represent a user's emotions and concerns, consider a user input reasoning task as "What's the latest server? Tell me about it quickly!" Assuming the user's emotion in this reasoning task is urgency and their concern is the latest server, the computing device can use "urgency" and "latest server" as paralinguistic information for this reasoning task.
[0084] It is understandable that the process by which computing devices acquire paralinguistic information contained in reasoning tasks varies depending on the specific reasoning task.
[0085] The following describes the process by which a computing device acquires paralinguistic information contained in a reasoning task, taking as examples 1, where the reasoning task includes audio information, and 2, where the reasoning task includes text information.
[0086] 1. The reasoning task includes audio information.
[0087] The computing device can first determine the acoustic features in the audio information, and then determine the para-language information based on the acoustic features in the audio information.
[0088] When the reasoning task involves audio information, the paralinguistic information may include stressed words from the audio information and / or emotional features contained within the audio information. Stressed words represent the user's focus, while emotional features represent the user's emotions.
[0089] The acoustic features of audio information can include physical features that characterize the objective attributes of the audio information, such as fundamental frequency, spectral characteristics, energy, zero-crossing rate, and duration. Acoustic features can also include perceptual features that characterize the user's subjective experience of the audio information, such as pitch and rhythmic features perceived by the user when listening to audio.
[0090] Among them, pitch features can include features such as fundamental frequency and formant frequency, while prosodic features can include features such as speech rate, pause patterns, and stress distribution.
[0091] When acoustic features include physical features, object features can be used to determine stressed words in audio information. Accordingly, the aforementioned computing device's determination of paralinguistic information based on acoustic features can be replaced by: determining stressed words based on physical features.
[0092] Taking object features as spectral features as an example, after receiving audio information, the computing device can use a preset algorithm to determine the spectral features of the audio information. Then, based on the spectral features, it can determine the probability of stressed words appearing in the audio of each time frame. The audio corresponding to multiple time frames where the probability of consecutive stressed words appearing is greater than a preset probability threshold is defined as stressed intervals. Subsequently, the computing device can use ASR (Advanced Sequencing Ratio) technology to extract stressed words from the stressed intervals. In this way, the computing device can obtain the stressed words contained in the audio information.
[0093] The embodiments of this application do not specifically limit the preset algorithm. For example, the preset algorithm can be MFCC, perceptual linear prediction (PLP), or linear-predictive cepstral coefficients (LPCC).
[0094] Through the above technical solution, the computing device can identify the stressed words in the audio information. Stressed words can usually represent the user's focus. In this way, the computing device can obtain the user's focus through stressed words, so that the computing device can adjust the inference results based on the user's focus to meet the user's needs, improve the user's satisfaction with the inference results, and enhance the user's experience.
[0095] To effectively improve the efficiency of computing devices in determining stressed words, in one optional implementation, a stressed word inference model can be deployed in the computing device, and this model is used to determine stressed words based on physical features. Based on this, the aforementioned determination of stressed words based on physical features can be replaced by: inputting audio information into the stressed word inference model to obtain the stressed words.
[0096] This application does not specifically limit the stress word inference model. For example, the stress word inference model can be a model built based on transformer networks, recursive neural networks (RNNs), or convolutional neural networks (CNNs).
[0097] Specifically, after receiving audio information, the computing device can input the audio information into the stress word inference model. The stress word inference model can determine the physical characteristics of the audio information and, based on the physical characteristics of the audio information, determine the stress words contained in the audio information.
[0098] Before using the stress word inference model, the computing device can also use pre-labeled audio data with stress words (hereinafter referred to as sample data) to train the initial model and obtain the stress word inference model.
[0099] Specifically, let's take a CNN-based stress word inference model as an example. The computing device can input sample data into an initial CNN-based model, which can predict stress words in the sample data. The computing device can then determine the difference between the predicted results and the pre-labeled stress words in the sample data (hereinafter referred to as the true results), and optimize the initial model based on this difference until the model converges, thus obtaining the stress word inference model.
[0100] In one example, the aforementioned computing device can determine the difference between the predicted result and the actual result using the cross-entropy loss function shown in Formula 1 below.
[0101]
[0102] Where y represents the actual result. The value represents the predicted result, and the loss represents the difference between the predicted result and the actual result.
[0103] This application does not specifically limit the method by which the computing device optimizes the initial model. For example, the computing device can optimize the initial model using the Adam algorithm or gradient descent.
[0104] Through the above technical solution, the computing device can quickly obtain the stressed words in the audio information through the stressed word inference model. In this way, the efficiency of the computing device in determining stressed words can be effectively improved, thereby effectively improving the efficiency of the computing device in reasoning for the inference task input by the user.
[0105] When acoustic features include perceptual features, the perceptual features can be used to determine the emotional features in the audio information. Accordingly, the above-mentioned computing device can determine the paralinguistic information based on acoustic features, which can be specifically implemented as: determining the emotional features in the audio information based on perceptual features.
[0106] Specifically, taking the perceptual features, including pitch features and prosodic features, as an example, after receiving audio information, the computing device can determine the pitch features and prosodic features of the audio information, and based on the pitch features and prosodic features of the audio information, determine the emotional features contained in the audio information.
[0107] This application does not specifically limit the emotional features corresponding to different perceptual features. For example, when the pitch feature is sharp and the rhythm feature includes fast speech rate and frequent pauses, the corresponding emotional feature can be anger. As another example, when the pitch feature is low and the rhythm feature is slow speech rate, the corresponding emotional feature can be surprise.
[0108] In one alternative implementation, the computing device may include Figure 5 The illustrated paralinguistic information extraction module is used to determine the paralinguistic information contained in the audio information. In one example, the paralinguistic information extraction module can be used to determine stressed words and / or emotional features contained in the audio information.
[0109] To effectively improve the efficiency of computing devices in determining emotional features, in one optional implementation, the computing device may be equipped with an emotion inference model, which is used to determine emotional features based on perceptual features. Based on this, the above-mentioned determination of emotional features in audio information based on perceptual features can be replaced by: inputting the audio information into the emotion inference model to obtain the emotional features.
[0110] This application does not specifically limit the emotion reasoning model. For example, the emotion reasoning model can be a model built based on transformer networks, RNNs, or CNNs.
[0111] Specifically, after receiving audio information, the computing device can input the audio information into the emotion reasoning model. The emotion reasoning model can determine the perceptual features of the audio information and, based on the perceptual features, determine the emotional features contained in the audio information.
[0112] Before using the sentiment inference model, the computing device can also train the initial model to obtain the sentiment inference model. The specific training process can be found in the description of training the stressed word inference model above, and will not be repeated here.
[0113] 2. The reasoning task includes textual information.
[0114] The computing device can first identify keywords in the text information, and then determine the paralinguistic information based on the keywords.
[0115] Keywords may include at least one of the following: words used to characterize entities (such as server, computer, etc.), words used to characterize user emotions (such as anxious, as soon as possible, hurry up, happy, etc.), or words whose frequency of occurrence is greater than a threshold (hereinafter referred to as frequency threshold).
[0116] This application does not specifically limit the keywords. For example, keywords can be punctuation marks, such as question marks and exclamation marks, or they can be modal particles, such as ya, ah, and ba.
[0117] Specifically, taking the inclusion of words with a frequency greater than a threshold as an example, the computing device can perform word segmentation on the inference task and remove stop words from the segmented inference task. Then, the computing device can determine the frequency of each word after stop word removal in the inference task. Subsequently, the computing device can use words with a frequency greater than a frequency threshold as keywords and analyze user sentiment and / or user focus based on these keywords, thus obtaining the paralinguistic information contained in the inference task.
[0118] Stop words refer to words that appear frequently but have low semantic value. They are usually function words in the language, such as articles, prepositions, conjunctions, and auxiliary verbs.
[0119] In one example, computing devices can use NLP technology to analyze users' emotions based on keywords.
[0120] To effectively improve the efficiency of word segmentation for inference tasks, in one optional implementation, a word segmentation model can be deployed in the computing device. This word segmentation model can perform word segmentation for inference tasks.
[0121] S302, process the reasoning task to obtain the first reasoning result.
[0122] This application does not impose a specific limitation on the order in which the computing device executes S301 and S302. For example, the computing device may execute S301 first and then S302. Alternatively, the computing device may execute S302 first and then S301. Yet another example is that the computing device may execute S301 and S302 simultaneously.
[0123] Specifically, the computing device stores multiple reference tasks and corresponding reference results for each reference task. Accordingly, the computing device can first determine the similarity between each reference task and the inference task, and then use the reference result corresponding to the reference task with the highest similarity to the inference task as the first inference result.
[0124] This application does not limit the storage format of the reference task and the reference result. For example, the reference task and the reference result can be stored in text form or in vector form.
[0125] In one example, the vector corresponding to the reference task (hereinafter referred to as the task vector) and the vector corresponding to the reference result (hereinafter referred to as the result vector) can be stored in a vector database (or local knowledge base).
[0126] For example, taking the storage of reference tasks and reference results as vectors, and the storage of task vectors and result vectors in a vector database as an example, the reference... Figure 4 As shown, after receiving a reasoning task input by the user, the computing device can first determine the type of the reasoning task. If the reasoning task is non-textual information (such as audio, video, or image information), the computing device can convert the reasoning task into textual information (such as...). Figure 4 (Data analysis shown). Afterwards, the computing device can convert the text information corresponding to the inference task into vectors (hereinafter referred to as inference vectors, such as...) using an embedding model. Figure 4 (The data processing is shown). Afterwards, the computing device can determine the task vector (e.g., ...) that is most similar to the inference vector from the vector database. Figure 4 The vector retrieval shown above converts the corresponding result vector into text (e.g., ...). Figure 4 (The vector transformation shown) yields the first inference result.
[0127] In one alternative implementation, the computing device may include Figure 5 The text conversion module shown is used to convert information of different modalities (such as audio information, video information, or images) into text information.
[0128] When the reasoning task includes audio information, the electronic device can convert the audio information into text information (hereinafter referred to as audio text) before processing the reasoning task and obtaining the first reasoning result. Accordingly, the electronic device can process the reasoning task and obtain the first reasoning result instead of processing the audio text and obtaining the first reasoning result.
[0129] This application does not specifically limit the method of converting audio information into text information. For example, it can be done through a speech recognition model built by a transformer network, RNN or CNN, or through ASR technology.
[0130] In order to effectively improve the efficiency of the computing device in obtaining the first inference result, in an optional implementation, the computing device may be equipped with an inference model. Accordingly, the computing device may input the inference task into the inference model, and the inference model may process the inference task and output the first inference result.
[0131] This application does not specifically limit the inference model. For example, the inference model can be an LLM built based on a transformer network, RNN or CNN, a multimodal question answering model, an embedding model, etc.
[0132] S303, based on the secondary language information, adjust the first reasoning result to obtain the second reasoning result.
[0133] S304, output the second reasoning result.
[0134] Specifically, the matching degree between the second inference result and the paralinguistic information is greater than that between the first inference result and the paralinguistic information. The matching degree between the inference result (such as the first inference result or the second inference result) and the paralinguistic information can be determined by the user's satisfaction with the inference result. That is, the matching degree between the second inference result and the paralinguistic information is greater than that between the first inference result and the paralinguistic information, which can also be described as: the user's satisfaction with the second inference result is greater than the user's satisfaction with the first inference result.
[0135] Specifically, the computing device can store response templates corresponding to different paralinguistic information. Specifically, after obtaining the paralinguistic information and the first inference result in the inference task, the user can determine a response template from their stored response templates that matches the paralinguistic information in the inference task, and use that response template to adjust the first inference result to obtain the second inference result.
[0136] For example, taking the reasoning task as "Please tell me as soon as possible why server A's performance is...", the first reasoning result is "Server A's performance is XXXXX", and the paralinguistic information in the reasoning task represents the user's emotion as urgency. After the computing device adjusts the first reasoning result using a response template corresponding to the urgency emotion, it can obtain the second reasoning result: "Sorry to keep you waiting, server A's performance is XXXXX". Then, the computing device can output "Sorry to keep you waiting, server A's performance is XXXXX".
[0137] Through the above technical solution, when a computing device performs reasoning on a user-input reasoning task, it can take into account information such as the user's emotions and / or concerns implicit in the reasoning task, and thus adjust the reasoning result based on the user's emotions and / or concerns. In this way, the reasoning result output by the computing device (such as the second reasoning result) can better match the user's expectations, effectively improving user satisfaction with the reasoning result output by the computing device and enhancing the user experience.
[0138] Furthermore, in this embodiment, the reasoning task input by the user can be multimodal data such as audio information, video information, image information, or text information, which can further improve the user experience.
[0139] To effectively improve the efficiency of adjusting the first inference result to obtain the second inference result, in one optional implementation, a reasoning model can be deployed in the computing device, which can adjust the first inference result. Accordingly, the above-mentioned S303 can be replaced by: determining the prompt word corresponding to the sub-language information, inputting the prompt word corresponding to the sub-language information and the first inference result into the inference model to obtain the second inference result.
[0140] The cue words are used to interact with the inference model so that the inference model can output content that matches the cue words.
[0141] Specifically, the computing device can store prompt words corresponding to different paralinguistic information. Specifically, after obtaining the paralinguistic information and the first inference result in the inference task, the user can determine the prompt word that matches the paralinguistic information in the inference task from their stored prompt words, and input the prompt word and the first inference result into the inference model. The inference model can then adjust the first inference result based on the prompt word to obtain a second inference result.
[0142] The above technical solution can not only effectively improve the efficiency of adjusting the first reasoning result, but also allow the reasoning model to fully consider the user's emotions and / or focus when generating the second reasoning content. In this way, the reasoning result output by the computing device (i.e. the second reasoning result) can better meet the user's expectations, effectively improve the user's satisfaction with the reasoning result output by the computing device, and enhance the user's experience.
[0143] To make the inference results output by the computing device more comprehensive and further improve user satisfaction with the inference results, in one optional implementation, the computing device may determine the entities included in the inference task (hereinafter referred to as the first entity) before processing the inference task to obtain the first inference result. Accordingly, the computing device processing the inference task to obtain the first inference result can be replaced by: determining the target reference task with the highest similarity to the inference task from multiple reference tasks, obtaining the entity information of the entity specifically associated with the first entity (hereinafter referred to as the second entity), and determining the first inference result based on the reference result corresponding to the target reference task and the entity information of the second entity.
[0144] An entity can be a concrete object, such as a person, a place, or an item, or it can be an abstract concept, such as an event or an idea.
[0145] This application does not specifically limit the entity information. For example, entity information may include at least one of the following: the entity's name, type, material, color, and other attribute information; the entity's location, distribution range, and other control information; and the entity's function, performance, and other information.
[0146] The process by which the computing device determines the target reference task can be referred to in S302, and will not be repeated here.
[0147] Specifically, the computing device can store the relationships between different entities, as well as the entity information of each entity. The computing device can obtain the first entity contained in the reasoning task through natural language processing technology, and determine the second entity that is related to the first entity based on its stored relationships. Then, the computing device can obtain the entity information of the second entity from its stored entity information.
[0148] After obtaining the entity information of the second entity and the reference result corresponding to the target reference task, the computing device can integrate the entity information of the second entity and the reference result corresponding to the target reference task to obtain the first inference result.
[0149] As can be seen from the above technical solution, the first reasoning result not only includes the result obtained after reasoning about the reasoning task, but also includes entity information of the second entity associated with the first entity in the reasoning task. In this way, the content of the first reasoning result can be richer and more comprehensive.
[0150] To accelerate the efficiency of the computing device in obtaining entity information of a second entity, in one optional implementation, a knowledge graph may be deployed in the computing device. The knowledge graph stores entities that are associated with different entities, as well as entity information for each entity. Accordingly, the aforementioned acquisition of entity information of the second entity by the computing device can be replaced by: determining the second entity from the knowledge graph and acquiring the entity information of the second entity.
[0151] A knowledge graph is a data structure used to represent different entities and the relationships between them. Specifically, a knowledge graph can use nodes to represent entities and edges between different nodes to represent the relationships between different entities.
[0152] In one example, a computing device can retrieve entity information of a second entity from a knowledge graph using the SPARQL query language.
[0153] In one alternative implementation, a retrieval-enhanced generation (RAG) system may be deployed in the computing device. The RAG system is connected to the knowledge graph, and the computing device can obtain entity information of a second entity from the knowledge graph through the RAG system.
[0154] In one alternative implementation, the computing device may include Figure 5 The entity extraction module shown is used to determine the second entity from the knowledge graph and obtain the entity information of the second entity.
[0155] To avoid errors in the reasoning task, such as typos or grammatical errors, that could lead to deviations in the first entity determined from the reasoning task, in one alternative implementation, the computing device may further correct the reasoning task before determining the first entity contained in the reasoning task, and determine the first entity from the corrected reasoning task.
[0156] This application does not set a specific precedent for modifying the reasoning task. For example, it can correct typos in the reasoning task, correct grammatical errors in the reasoning task, or delete characters in the reasoning task.
[0157] To effectively improve the efficiency of correcting inference tasks, in one optional implementation, a correction model can be deployed in the computing device. Accordingly, the above-mentioned correction of the inference task by the computing device can be replaced by: the computing device inputting the inference task into the correction model to obtain the corrected inference task.
[0158] By correcting the reasoning task in the above way, not only can the accuracy of extracting the first entity be effectively improved, but the accuracy of subsequent computing devices in determining the target reference task can also be effectively improved.
[0159] To provide a more detailed and comprehensive overview Figure 3 The reasoning method shown will be introduced below. Taking the reasoning task containing audio information and the paralinguistic information including accented words and emotional features in the audio information as an example, the reasoning amplification provided by the embodiments of this application will be introduced in conjunction with the above embodiments. Figure 6 A flowchart illustrating another reasoning method provided in this application embodiment is shown below. Figure 6 As shown, the method includes the following steps S601-S612.
[0160] S601 receives audio information input by the user.
[0161] S602 performs preprocessing on audio information.
[0162] Preprocessing audio information can include denoising and / or enhancing the audio information.
[0163] This application does not specifically limit the method of the computing device for denoising audio information. For example, the computing device may use methods such as spectral subtraction, low-pass filtering, or high-pass filtering to denoise audio information.
[0164] Taking spectral subtraction as an example, the computing device can perform noise reduction on audio information based on Fourier transform using the following formula 2.
[0165] x(t)=s(t)+n(t) (Formula 2)
[0166] Where x(t) represents the audio information containing noise, s(t) represents the original audio information, i.e., the audio information without noise, and n(t) represents the noise (also known as additive noise).
[0167] This application does not specifically limit the method of the computing device for enhancing audio information. For example, the computing device may use equalization processing or filtering processing to enhance audio information. For specific processes, please refer to relevant technologies, which will not be elaborated here.
[0168] It should be noted that step S602 is optional.
[0169] S603, determine the emotional features in the audio information.
[0170] S604, Identify stressed words in audio information.
[0171] The processes of S603 and S604 can be referred to the description in S301 above, and will not be repeated here.
[0172] For example, taking the user-input audio information as "Does your company have any of the latest servers? Are there any particularly good devices? Please tell me about them!" with "latest" as the stressed word, the computing device can determine that the emotional feature contained in the audio information is "anxious" and the stressed word is "latest".
[0173] S605 converts audio information into text information to obtain audio text.
[0174] S606, corrects audio text.
[0175] The process of S606 can be referred to the above description of the modification of the reasoning task, and will not be repeated here.
[0176] For example, continuing with the user-input audio message "Does your company have any of the latest servers? Are there any particularly good devices? Please tell me about them!", the corrected audio text could be "Please tell me about your company's latest servers."
[0177] S607, determine the reference result corresponding to the target reference task with the highest similarity to the corrected audio text.
[0178] S608, Identify the first entity from the corrected audio text.
[0179] For example, continuing with the revised audio text "Give me an introduction to your company's latest server," the computing device can use "server" as the first entity.
[0180] S609, Based on the first entity, determine the entity information of the second entity from the knowledge graph.
[0181] S610, based on the reference result and the entity information of the second entity, determine the first reasoning result.
[0182] For example, continuing with the revised audio text "Please introduce your company's latest server to me," assuming the company's latest server is server A, and server B is associated with server A, the computing device can use "server B" as a second entity. Information related to "server B" will then be used as the entity information of this second entity. Taking the server entity information as including the server type (server A is type A, server B is type B) as an example, the first inference result could be: Server A, model A; Server B, model B.
[0183] S611, determine the response template corresponding to emotional characteristics and stressed words.
[0184] S612, the second inference result is adjusted using the response template corresponding to the emotional features and stressed words to obtain the second inference result.
[0185] For example, assuming the emotional characteristic is "anxious" and the stressed word is "latest," the first inference result is: Server A, model A; Server B, model B. The second inference result, after the computing device adjusts it using a response template, could be: "I apologize for keeping you waiting. Please allow me to introduce them to you one by one. Our company's latest server is Server A, model A. Also related is Server B, model B."
[0186] S612, output the second reasoning result.
[0187] As described above, the second inference result output by the computing device can not only take into account the user's emotions (anxiety) and / or the user's focus (such as the latest information), but also include entity information related to the entities in the inference task. This not only makes the inference result output by the computing device more aligned with the user's expectations, but also makes the output more comprehensive, thus effectively improving user satisfaction with the inference result output by the computing device and enhancing the user experience.
[0188] It should be noted that the reasoning method provided in this application embodiment can be applied in different scenarios, such as mental health support platforms and customer support centers.
[0189] In a customer support center scenario, the customer support center can efficiently and accurately answer user questions using the reasoning method provided in this application embodiment. It can not only provide comprehensive reasoning results through knowledge graphs, but also optimize the reasoning results through user emotional characteristics, thereby improving customer satisfaction.
[0190] In the context of a mental health support platform, users can describe their psychological distress through language, text, or video. The mental health support platform can use the reasoning methods provided in this application to understand the user's emotions and concerns, so as to provide comprehensive advice to the user and help the user improve their mental health.
[0191] The foregoing primarily describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the aforementioned functions, the inference device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0192] This application embodiment can, according to the above method, exemplarily divide the inference device into functional modules. For example, the inference device may include functional modules corresponding to each functional division, or two or more functions may be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0193] For example, Figure 7 A possible structural schematic diagram of the inference device involved in the above embodiments is shown. The device includes: an acquisition unit 701, a processing unit 702, an adjustment unit 703, and an output unit 704. The acquisition unit 701 is used to acquire paralinguistic information contained in the inference task in response to user input. The processing unit 702 is used to process the inference task to obtain a first inference result. The adjustment unit 703 is used to adjust the first inference result based on the paralinguistic information to obtain a second inference result. The output unit 704 is used to output the second inference result.
[0194] The second inference result matches the paralinguistic information more closely than the first inference result. Paralinguistic information is used to represent the user's emotions and / or focus.
[0195] Optionally, the reasoning task may include audio information. Accordingly, the acquisition unit 701 described above can specifically be used to: determine paralinguistic information based on the acoustic features in the audio information.
[0196] Optionally, the aforementioned acoustic features may include physical features, which refer to the objective attributes of audio information. Paralinguistic information includes stressed words in the audio information. Accordingly, the aforementioned acquisition unit 701 can specifically be used to: determine stressed words based on physical features.
[0197] Optionally, a stress word inference model can be deployed in the computing device. This model is used to determine stress based on physical features. Accordingly, the acquisition unit 701 described above can specifically be used to: input audio information into the stress word inference model to obtain the stress word.
[0198] Optionally, the aforementioned acoustic features may include perceptual features, which characterize the user's subjective feelings about the audio information, and paralinguistic information includes emotional features in the audio information. Accordingly, the aforementioned acquisition unit 701 may specifically be used to: determine emotional features based on perceptual features.
[0199] Optionally, the adjustment unit 703 can be specifically used to: determine a response template that matches the sub-language information, adjust the first inference result using the response template, and obtain the second inference result.
[0200] Optionally, the computing device may store multiple reference tasks and reference results corresponding to each reference task. The processing unit 702 described above can also be used to determine a first entity contained in the inference task. Accordingly, the processing unit 702 can specifically be used to: determine the target reference task with the highest similarity to the inference task from multiple reference tasks, obtain entity information of the second entity, and determine a first inference result based on the reference result corresponding to the target reference task and the entity information of the second entity.
[0201] The second entity refers to an entity that has a specific relationship with the first entity.
[0202] Optionally, a knowledge graph may be deployed in the computing device. The knowledge graph stores entities that are associated with different entities, as well as entity information for each entity. Accordingly, the aforementioned acquisition unit 701 can be specifically used to: determine a second entity from the knowledge graph and acquire the entity information of the second entity.
[0203] Optionally, before determining the first entity included in the reasoning task, the computing device may further modify the reasoning task. Accordingly, the processing unit 702 may be specifically used to: determine the first entity from the modified reasoning task.
[0204] Optionally, the above reasoning task may include text information. Accordingly, the above acquisition unit 701 may be used to: determine keywords in the text information, and determine paralinguistic information based on the keywords.
[0205] Keywords may include at least one of the following: words used to represent entities, words used to represent user emotions, or words that appear more frequently than a threshold.
[0206] For a detailed description of the above-mentioned optional methods, please refer to the foregoing method embodiments, which will not be repeated here. Furthermore, the explanation of any of the inference devices provided above and the description of their beneficial effects can be found in the corresponding method embodiments described above, which will not be repeated here.
[0207] This application also provides a computing device, which includes a processor and a memory. The processor is connected to the memory, and the memory stores computer execution instructions. When the processor executes the computer execution instructions, it implements the data processing method described in the above embodiments. This application does not limit the specific form of the computing device. For example, the computing device can be a terminal device or a network device. The terminal device can be referred to as: terminal, user equipment (UE), terminal device, access terminal, user unit, user station, mobile station, remote station, remote terminal, mobile device, user terminal, wireless communication device, user agent, or user equipment, etc. The terminal device can specifically be a mobile phone, augmented reality (AR) device, virtual reality (VR) device, tablet computer, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc. The network device can specifically be a server, etc. The server can be a single physical or logical server, or two or more physical or logical servers sharing different responsibilities and cooperating to achieve the various functions of the server.
[0208] This application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods executed by any of the computing devices described above.
[0209] For explanations of the relevant content and descriptions of the beneficial effects in any of the computer-readable storage media provided above, please refer to the corresponding embodiments described above, which will not be repeated here.
[0210] This application also provides a chip. The chip integrates a control circuit for implementing the functions of the aforementioned computing device and one or more ports. Optionally, the functions supported by the chip can be referred to above, and will not be repeated here. Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium. The aforementioned storage medium can be a read-only memory, random access memory, etc. The aforementioned processing unit or processor can be a central processing unit, a general-purpose processor, an application-specific integrated circuit (ASIC), a microprocessor (digital signal processor, DSP), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.
[0211] This application also provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform any of the methods described in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or may include one or more data storage devices such as servers or data centers that can be integrated with the medium. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., SSD), etc.
[0212] It should be noted that the devices for storing computer instructions or computer programs provided in the embodiments of this application, such as but not limited to the memory, computer-readable storage medium and communication chip, are all non-transitory.
[0213] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs).
[0214] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, the disclosure, and the appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0215] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.
Claims
1. A reasoning method, characterized in that, Applied to a computing device, the method includes: In response to a reasoning task input by a user, the system obtains paralinguistic information contained in the reasoning task, which is used to characterize the user's emotions and / or focus. The reasoning task is processed to obtain the first reasoning result; Based on the paralinguistic information, the first inference result is adjusted to obtain a second inference result, wherein the matching degree between the second inference result and the paralinguistic information is greater than the matching degree between the first inference result and the paralinguistic information. Output the second reasoning result.
2. The method according to claim 1, characterized in that, The reasoning task includes audio information, and the acquisition of paralinguistic information contained in the reasoning task includes: The secondary language information is determined based on the acoustic features in the audio information.
3. The method according to claim 2, characterized in that, The acoustic features include physical features, which refer to the objective attributes of the audio information; the paralinguistic information includes stressed words in the audio information. The determination of the secondary language information based on the acoustic features in the audio information includes: Based on the physical characteristics, the stressed word is determined.
4. The method according to claim 3, characterized in that, The computing device is equipped with a stress word inference model, which is used to determine the stress word based on the physical features. The determination of the stressed word based on the physical characteristics includes: The audio information is input into the stress word inference model to obtain the stress word.
5. The method according to any one of claims 2-4, characterized in that, The acoustic features include perceptual features, which are used to characterize the user's subjective feelings about the audio information; the paralinguistic information includes emotional features in the audio information. The determination of the paralinguistic information based on the acoustic features includes: Based on the perceptual features, the emotional features are determined.
6. The method according to any one of claims 1-4, characterized in that, The step of adjusting the first inference result based on the secondary language information to obtain the second inference result includes: Determine a response template that matches the aforementioned secondary language information; The first reasoning result is adjusted using the aforementioned response template to obtain the second reasoning result.
7. The method according to any one of claims 1-4, characterized in that, The computing device stores multiple reference tasks and reference results corresponding to each reference task; Before processing the reasoning task to obtain the first reasoning result, the method further includes: Identify the first entity contained in the reasoning task; The process of processing the reasoning task to obtain the first reasoning result includes: From the plurality of reference tasks, determine the target reference task with the highest similarity to the reasoning task; Obtain the entity information of the second entity; the second entity refers to an entity that is specifically associated with the first entity; Based on the reference result corresponding to the target reference task and the entity information of the second entity, the first reasoning result is determined.
8. The method according to claim 7, characterized in that, The computing device is equipped with a knowledge graph, which stores entities that are associated with different entities, as well as entity information for each entity. The step of obtaining the entity information of the second entity includes: The second entity is determined from the knowledge graph, and the entity information of the second entity is obtained.
9. The method according to claim 7, characterized in that, Before determining the first entity included in the reasoning task, the method further includes: The reasoning task is modified accordingly; Determining the first entity included in the reasoning task includes: The first entity is determined from the revised reasoning task.
10. The method according to claim 1, characterized in that, The reasoning task includes text information, and obtaining the paralinguistic information contained in the reasoning task includes: The keywords in the text information are determined, and the keywords include at least one of the following: words used to represent entities, words used to represent user emotions, or words that appear more frequently than a threshold. Based on the keywords, the secondary language information is determined.
11. A computing device, characterized in that, include: Processor and memory; The processor is connected to a memory for storing computer execution instructions, and the processor executes the computer execution instructions stored in the memory to enable the computing device to implement the method as described in any one of claims 1-10.