Voice Interaction Method, Apparatus, Device, and Storage Medium
By obtaining voice information to determine audio characteristics, automatically selecting the target service mode and outputting resources, the problem of users needing to manually select modes in the prior art is solved, and convenient voice interaction and efficient resource management are achieved.
Patent Information
- Application Number
- CN202111221780.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-10-20
AI Technical Summary
Existing smart voice devices require users to manually select user mode or carry mode information in voice commands, which increases user operations and reduces user experience.
By obtaining voice information, determining audio characteristics, automatically determining the target service mode based on the audio characteristics, and correlating the corresponding resource collection, outputting the target resources, and reducing user operations.
It improves the operation convenience of the voice interaction process, enhances the user experience, avoids inappropriate resource output, and improves the accuracy and efficiency of resource collection construction.
Smart Images

Figure CN113963687B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to speech technologies, and specifically to a speech interaction method, apparatus, device, and storage medium. Background Art
[0002] With the continuous development of artificial intelligence technologies, intelligent speech devices (such as intelligent speakers) have emerged, providing convenience for users' lives. For example, intelligent speech devices can provide feedback on corresponding resources according to users' speech commands.
[0003] In the prior art, when using an intelligent speech device, a user needs to manually select a user mode or carry mode information in a speech command to select a user mode, which requires the user's cooperation in operation, increasing the user's operations and reducing the user's experience. Summary of the Invention
[0004] The present disclosure provides a speech interaction method, apparatus, device, and storage medium.
[0005] According to one aspect of the present disclosure, there is provided a speech interaction method, including:
[0006] Obtaining speech information;
[0007] Determining an audio feature according to the speech information;
[0008] Determining a target service mode according to the audio feature;
[0009] Determining a target resource for output according to a resource set associated with the target service mode.
[0010] According to another aspect of the present disclosure, there is also provided an electronic device, including:
[0011] At least one processor; and
[0012] A memory communicatively connected to the at least one processor; wherein,
[0013] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any speech interaction method provided by the embodiments of the present disclosure.
[0014] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any speech interaction method provided by the embodiments of the present disclosure.
[0015] According to the technology of the present disclosure, the operation convenience is improved.
[0016] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood from the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0018] Figure 1A is a schematic diagram of a voice interaction method provided according to an embodiment of the present disclosure;
[0019] Figure 1B is a schematic diagram of a voice interaction system provided according to an embodiment of the present disclosure;
[0020] Figure 1C is a schematic diagram of another voice interaction system provided according to an embodiment of the present disclosure;
[0021] Figure 2 is a schematic diagram of another voice interaction method provided according to an embodiment of the present disclosure;
[0022] Figure 3 is a schematic diagram of another voice interaction method provided according to an embodiment of the present disclosure;
[0023] Figure 4 is a schematic diagram of another voice interaction method provided according to an embodiment of the present disclosure;
[0024] Figure 5 is a structural diagram of a voice interaction device provided according to an embodiment of the present disclosure;
[0025] Figure 6 is a block diagram of an electronic device for implementing the voice interaction method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding and should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0027] The various voice interaction methods and voice interaction devices provided by the embodiments of the present disclosure are applicable to application scenarios for voice interaction with intelligent voice devices. The various voice interaction methods provided by the embodiments of the present disclosure can be executed by a voice interaction device, which can be implemented by software and / or hardware and is specifically configured in an electronic device. The electronic device can be a voice interaction device or other computing devices associated with the voice interaction device. Exemplarily, the voice interaction device can be a mobile phone or a tablet, etc. In particular, the voice interaction device can be a smart speaker.
[0028] For ease of understanding, the various voice interaction methods provided by the present disclosure will be described in detail below.
[0029] See Figure 1A A voice interaction method shown in the figure includes:
[0030] S110. Obtain voice information.
[0031] Among them, the voice information can be obtained by an intelligent voice device. For example, it can be a smart speaker device, or a mobile phone, a tablet or a notebook with voice interaction function, etc. It can also be collected by the intelligent voice device and transmitted to other computing devices associated with the intelligent voice device, such as a cloud server, for use. The embodiments of the present disclosure do not impose any limitations on the operating system of the cloud server. For example, the DuerOS operating system can be adopted.
[0032] S120. Determine an audio feature according to the voice information.
[0033] Exemplarily, the voiceprint information in the voice information can be extracted to obtain the audio feature. Among them, the audio feature can include voiceprint information, carrying frequency domain characteristic information and / or time domain characteristic information. For example, the voiceprint information in the voice information can be extracted through voiceprint feature extraction technology to obtain the audio feature. In an alternative embodiment, the voice information can be input into a trained voiceprint extraction model, and the audio feature can be determined according to the output result. Among them, the voiceprint extraction model can be implemented based on a machine learning model or a deep learning model, and the present disclosure does not impose any limitations on this.
[0034] S130. Determine a target service mode according to the audio feature.
[0035] Exemplarily, different service modes are set according to different target audiences for use. In a specific implementation, the service mode can be set as an elderly mode, a children's mode, a youth mode, a middle-aged mode, etc. according to the age attribute of the target audience; different service modes for different regions can also be set according to the region; different function modes can also be set according to different required functions, such as an entertainment mode, a finance mode, etc.; male and female modes can also be set according to gender. The embodiments of the present disclosure can be set or adjusted by technicians or operators according to actual needs. Among them, the target service mode can be understood as the service mode that matches the initiator of the voice information.
[0036] In an alternative embodiment, the audio feature can be input into a trained classification model, and according to the classification result, the service mode corresponding to the audio feature is determined. Among them, the classification model can be obtained by training a pre-constructed machine learning model or a deep learning model with a large number of sample audio features and corresponding service mode labels. Among them, the specific network structure of the machine learning model or the deep learning model is not limited in the present disclosure.
[0037] In another alternative embodiment, for example, in the home application scenario of an intelligent voice device, the operator of the intelligent voice device is relatively fixed. Therefore, the audio features of different operators can also be pre-stored, and the service mode of the corresponding operator is set. Correspondingly, the audio feature determined according to the voice information is matched with the pre-stored audio features, and the service mode corresponding to the matched audio feature is used as the target service mode.
[0038] S140. Determine a target resource for output according to the resource set associated with the target service mode.
[0039] Among them, the resource set can include at least one of an audio resource set, a video resource set, a picture resource set, a text resource set, etc. The target resource can be an element in the resource set corresponding to the target service mode.
[0040] Exemplarily, resource labels corresponding to each resource can be determined, and resource sets associated with different service modes are constructed through the resource labels. For example, the resource label corresponding to a certain resource "Movie A" can include a "suspense" label. It should be noted that any resource can correspond to at least one type of label. For example, the resource labels corresponding to a certain picture resource can include a "colorful system" label and an "animal" label, etc. Exemplarily, clustering analysis can be performed on each resource label through a clustering algorithm, and different resource sets are constructed according to the clustering result. For example, the clustering algorithm can be a k-means clustering algorithm (k-means clustering algorithm, K-means).
[0041] In an alternative embodiment, the correspondence between the service mode and the resource set can be bound by a technician according to actual requirements, so as to achieve resource constraints for different audience groups by restricting the resource sets of different service modes.
[0042] In another alternative embodiment, the correspondence between different service modes and resource sets can also be determined in an automated manner, thereby improving the determination efficiency of the above correspondence and reducing labor costs. Exemplarily, for a certain service mode, candidate resources can be selected from the original resources according to the associated keywords of the service mode; the candidate resources are added to the resource set of the service mode.
[0043] Among them, the original resources can be understood as the full amount of resources in the original resource library that can be output when using an intelligent voice device.
[0044] It can be understood that the above technical solution constructs the resource set of the corresponding service mode automatically by introducing the associated keywords of the service mode, avoiding the poor accuracy of the determination result of candidate resources caused by the differences in manual construction and human fatigue, and then improving the accuracy of the construction result of the resource set corresponding to different service modes. At the same time, by constructing the resource sets of different service modes, the resources that can be output by different service modes are restricted, avoiding providing inappropriate resource content to the audience group under the service mode, and improving the usage experience of the voice information initiator. At the same time, by the way of automatically constructing the resource set, the construction efficiency of the resource set is also improved, and the investment in labor costs is reduced.
[0045] Among them, the associated keywords of the service mode are used to represent the resource labels corresponding to the resources allowed or prohibited from being presented under the service mode. For example, the associated keywords can include permission keywords, corresponding to the resource labels of the resources allowed to be presented; another example is that the associated keywords can include taboo keywords, corresponding to the resource labels of the resources prohibited from being presented.
[0046] Optionally, for a certain service mode, the original resources marked with the taboo keywords of the service mode can be excluded from the original resources to obtain candidate resources; the candidate resources are added to the resource set of the service mode.
[0047] Specifically, for a certain service mode, the resource collection of the service mode can be pre-set to include all original resources, and then the original resources in the resource collection that are marked with the taboo keywords of the service mode are identified, and the resource collection after elimination is used as the resource collection of the service mode. Among them, the taboo keywords of the service mode can be set or adjusted by technical personnel according to needs or experience. For example, in the children's mode, keywords such as "violence" and "pornography" can be used as taboo keywords; accordingly, the original resources marked with resource tags such as "violence" or "pornography" are eliminated from the resource collection of the children's mode.
[0048] Alternatively, for a certain service model, original resources marked with the permission keyword of the service model may be selected as candidate resources; and the candidate resources may be added to the resource set of the service model.
[0049] Specifically, for a certain service mode, the resource set of the service mode can be pre-set to an empty set, and then the original resources marked with the permitted keywords of the service mode are determined as candidate resources; each candidate resource is added to the resource set of the service mode. Among them, the permitted keywords of the service mode can be set or adjusted by technical personnel according to needs or experience. For example, in the children's mode, keywords such as "comedy", "food" and "color" can be used as permitted keywords; accordingly, original resources marked with resource tags such as "comedy", "food" and "color" are selected from the original resources and added to the resource set of the children's mode.
[0050] It can be understood that the above scheme determines the candidate resources under a certain service model by elimination and / or selection, and then constructs the resource set of the service model, which enriches the construction method of the resource set and lays the foundation for determining the resource set of the target service model.
[0051] Exemplarily, determining the target resource based on the resource set associated with the target service mode can be: selecting at least one set element from the resource set associated with the target service mode as the target resource; and controlling the intelligent voice device to output the target resource to the initiator of the voice information.
[0052] Optionally, selecting at least one set element from the resource set associated with the target service mode as the target resource may be: randomly selecting at least one set element from the resource set associated with the target service mode as the target resource.
[0053] Alternatively, optionally, at least one set element is selected from the resource set associated with the target service mode as the target resource, which may be: at least one set element is selected from the resource set associated with the target service mode according to the initiation time or acquisition time of the voice information as the target resource. For example, the resource set stores the playing music of broadcast gymnastics and the background music for running; during the broadcast gymnastics playing period, the playing music of broadcast gymnastics is used as the target resource; during the running period, the background music for running is used as the target resource.
[0054] In the embodiments of the present disclosure, by introducing audio features to determine the target service mode, and restricting the content resources available for output through the resource set of the target service mode, the discomfort brought to the initiator of the voice information by the output of content resources in other service modes is avoided. In addition, in the embodiments of the present disclosure, the audio features are directly determined according to the voice information, and then the target service mode is automatically determined without manually inputting the target service mode, reducing the user operation, thereby improving the operation convenience of the voice interaction process and enhancing the user experience.
[0055] It should be noted that the execution subject of the voice interaction method in the embodiments of the present disclosure may be the intelligent voice device itself and / or other computing devices associated with the intelligent voice device to reduce the requirements for the computing power of the intelligent voice device. In an alternative embodiment, the voice interaction method may also be executed interactively by the intelligent voice device and at least one other computing device to achieve balanced distribution of computing resources.
[0056] See Figure 1B A voice interaction system shown in the figure includes an intelligent voice device 10 and a cloud server 20. Among them, the intelligent voice device 10 and the cloud server 20 are communicatively connected. The present disclosure does not make any limitation on the specific communication method and / or communication network.
[0057] Among them, the intelligent voice device 10 acquires voice information and sends the voice information to the cloud server 20. The cloud server 20 determines the audio features according to the voice information, and determines the target service mode according to the audio features; determines the target resource according to the resource set associated with the target service mode, and feeds back the target resource to the intelligent voice device 10. Among them, the determination operation of the target resource can refer to the description of other embodiments, and the present disclosure will not elaborate herein.
[0058] Exemplarily, the cloud server 20 can also cooperate with other platforms to complete related operations. For example, other platforms may include a skill platform and / or a resource platform. Specifically, see Figure 1CThe architecture diagram of the voice interaction system shown. The voice interaction system may include an intelligent voice device 10, a cloud server 20, a skill platform 30, and a resource platform 40. Among them, the skill platform 30 may include at least one of a feature recognition module, a mode configuration module, a dialogue management module, a label grouping module, etc.; the resource platform 40 may include different types of resources such as audio-visual resources, picture resources, text resources, and skill resources.
[0059] Specifically, the intelligent voice device 10 obtains voice information and sends the voice information to the cloud server 20; the cloud server 20 processes the voice information to determine the audio features; the cloud server 20 sends the audio features to the skill platform 30; the feature recognition module in the skill platform 30 determines the target service mode through the mode configuration module according to the audio features; the skill platform 30 may determine the resource set associated with the target service mode in the resource platform 40; the skill platform 30 may also determine the target resource according to the resource set associated with the target service mode and feedback the target resource to the intelligent voice device 10 through the cloud server 20. The label grouping module in the skill platform 30 may group the resource labels of different content resources in the resource platform 40 and determine the corresponding relationship between different groups and service modes through the mode configuration module, so as to realize the construction of the association relationship between the service mode and the resource set.
[0060] In an alternative embodiment, the skill platform 30 and / or the resource platform 40 may be integrally provided in the cloud server 20.
[0061] It can be understood that voice information is obtained through the intelligent voice device and sent to the cloud server for the determination of the target resource. Since the intelligent voice device only transmits voice information to the cloud server, the waste of bandwidth resources caused by the transmission of irrelevant data is reduced. At the same time, the determination process of the target resource is implemented in the cloud server, reducing the data operation amount of the intelligent voice device and the requirement for the data processing ability of the intelligent voice device, thereby reducing the hardware cost investment of the intelligent voice device.
[0062] Based on the above technical solutions, the present disclosure also provides an alternative embodiment. In this alternative embodiment, the determination operation of the target service mode is optimized and improved. For the parts not detailed in this embodiment, reference may be made to the description of the foregoing embodiments, which will not be elaborated herein.
[0063] See Figure 2 A voice interaction method shown, including:
[0064] S210. Obtain voice information.
[0065] S220. Determine audio features according to the voice information.
[0066] S230. Determine age information according to the audio features.
[0067] Among them, the age information may include an age value or an age range. Exemplarily, the audio features can be input into a trained age recognition model to obtain the age information. Among them, the age recognition model can be obtained by training a pre-constructed machine learning model or deep learning model based on a large number of sample audio features and age information labels. Among them, the sample audio features can be obtained by extracting voiceprint information from the sample voice information, and the age information labels can be manually marked or obtained by other existing methods. The present disclosure does not make any limitation on the specific network structure of the age recognition model.
[0068] In an optional embodiment, while determining the age information according to the audio features, the gender information can also be determined in association. For example, during the training process of the age recognition model, gender labels can be added, so that the trained age recognition model also has the ability to recognize gender. For example, the gender label of a female is set to 0, and the gender label of a male is set to 1. Of course, the gender labels can also be set to other different values respectively, and the present disclosure does not make any limitation on this.
[0069] S240. Determine the target service mode according to the age information.
[0070] Optionally, the age mode correspondence between different age information and service modes can be preset in advance; correspondingly, according to this correspondence, the target service mode matching the age information determined by the audio features is determined from each service mode.
[0071] In a specific implementation manner, the age mode correspondence can be the correspondence between different age groups and the corresponding service modes. For example, 0 - 12 years old corresponds to the children's mode; 55 years old and above corresponds to the elderly mode. It can be understood that by setting the correspondence between different service modes and the corresponding age groups, general scenarios can be adapted, such as home use scenarios or cinema use scenarios, etc. Taking the age information including an age value as an example, if the age value in the age information is 6 years old, the corresponding target service mode can be the children's mode; if the age value in the age information is 75 years old, the corresponding target service mode can be the elderly mode. Taking the age information including an age range as an example, if the age range in the age information is 7 - 9 years old, the corresponding target service mode can be the children's mode; if the age value in the age information is 70 - 75 years old, the corresponding target service mode can be the elderly mode.
[0072] In another specific implementation, the age pattern correspondence can be the correspondence between different age values and corresponding service modes. For example, 4 years old corresponds to the children's mode, 16 years old corresponds to the youth mode, 50 years old corresponds to the middle-aged mode, 70 years old corresponds to the elderly mode, etc. It can be understood that by setting the correspondence between different age values and corresponding service modes, application scenarios with relatively fixed user groups can be adapted, such as in the home use scenario.
[0073] Alternatively, optionally, the audio can be classified according to the audio features, and according to the audio classification result, the target service mode corresponding to each type of audio can be determined. Exemplarily, the audio features can be input into a trained audio classification model to obtain the audio classification result. Among them, the audio categories at least include children's category and elderly category, and can also include other categories, such as youth category. The target service mode is determined according to the audio category. For example, the target service mode corresponding to children's category audio is the children's mode. Among them, the audio classification model can be obtained by training a pre-constructed machine learning model or deep learning model based on a large number of sample audio features and corresponding audio classification labels. The present disclosure does not make any limitation on the specific network structure of the audio classification model.
[0074] Since the service modes expected to be used by the speech information originators of the same age may not be the same, and there may be certain errors in the age information determination results, there will be some controversy in the selected target service mode, which may affect the matching degree between the target service mode and the speech information originator. For example, some 12-year-old users expect to use the children's mode, while some 12-year-old users expect to use the youth mode. Another example is that the determined age information is 12 years old, while the actual age of the originator may be between 10 and 15 years old, and the children's mode can be applied to 10 - 12 years old, and the youth mode can be applied to 13 - 15 years old.
[0075] In order to further improve the matching degree between the determined result of the target service mode and the speech information originator, in an alternative embodiment, the association relationship between different age intervals and service modes can be preset in advance. Correspondingly, determining the target service mode according to the age information can include: determining the confidence type of the age information according to the age information and the adjacent age intervals of the age information; determining the target service mode from the service modes corresponding to the adjacent age intervals according to the confidence type.
[0076] Exemplarily, the service mode can be selected from the service modes corresponding to adjacent age intervals according to the confidence type as the target service mode; or, according to the confidence type, the target age interval can be selected from adjacent age intervals, and the service mode corresponding to the target age interval can be used as the target service mode. It can be understood that the determination of the target service mode is carried out by direct selection or indirect selection, enriching the diversity of the determination methods of the target service mode. For example, the service mode corresponding to the age interval of 1-12 years old is the children's mode; the service mode corresponding to the age interval of 13-18 years old is the adolescent mode; the service mode corresponding to the age interval greater than 69 years old is the elderly mode, etc.
[0077] Among them, the confidence type of the age information is used to characterize the credibility of the age information and / or the credibility of directly determining the corresponding service mode according to the age information. Exemplarily, the confidence type can include two categories: high confidence type and low confidence type. Among them, the adjacent age interval is used to characterize the age interval associated with the age information. For example, it can include the age interval to which the age value or age range in the age information belongs, and / or the adjacent age intervals of the age interval to which the age value or age range in the age information belongs. The adjacent age intervals can be two, such as the left neighbor or the right neighbor. Of course, the minimum age difference value between the age value or age range in the age information and the left and right adjacent age intervals can also be determined; the age interval with a smaller age difference is selected as the adjacent age interval.
[0078] In a specific implementation manner, the confidence type can be determined through a confidence interval. The confidence interval can include a high confidence interval and a low confidence interval. Among them, the low confidence interval can be set as a preset marginal sub-interval in the age interval; the high confidence interval can be set as a preset central sub-interval in the age interval; the preset marginal sub-interval is the complementary sub-interval of the preset central sub-interval in the corresponding age interval.
[0079] Among them, the preset edge sub-interval can be set in advance by those skilled in the relevant art. For example, if the age interval corresponding to the children's mode is 1 - 12 years old, the preset edge sub-interval can be set to 10 - 12 years old, and the preset central sub-interval can be 1 - 9 years old. Correspondingly, if the age value in the age information (such as 11 years old) corresponds to the preset edge sub-interval, it is considered that this age information belongs to the low confidence interval; if the age value in the age information (such as 8 years old) corresponds to the preset central sub-interval, it is considered that this age information belongs to the high confidence interval. If the age interval corresponding to the middle-aged mode is 46 - 69 years old, the preset edge sub-interval can be set to include 46 - 49 years old and 61 - 69 years old, and the preset central sub-interval can be set to 50 - 60 years old. Correspondingly, if the age value in the age information (such as 47, 62 years old) corresponds to the preset edge sub-interval, it is considered that this age information belongs to the low confidence interval; if the age value in the age information (such as 55 years old) corresponds to the preset central sub-interval, it is considered that this age information belongs to the high confidence interval.
[0080] Exemplarily, the adjacent age intervals of the age information, that is, the age interval to which the age value belongs and the adjacent age intervals of the age interval to which the age value belongs. For example, if the obtained age value is 11 years old, the age interval to which this age value belongs is 1 - 12 years old, and the adjacent age interval is 13 - 18 years old. Therefore, the adjacent age intervals corresponding to this age value are 1 - 12 years old and 13 - 18 years old. Another example, if the age interval corresponding to the youth mode is 14 - 17 years old (the corresponding preset central sub-interval is 15 - 16 years old), the age interval corresponding to the young people is 18 - 45 years old (the corresponding preset central sub-interval is 20 - 40 years old), and the age interval corresponding to the middle-aged is 46 - 69 years old (the corresponding preset central sub-interval is 50 - 60 years old). If the age value in the age information is 47 years old, it is determined that the age interval to which it belongs is 46 - 69 years old, and the adjacent age interval is 18 - 45 years old. Therefore, the adjacent age intervals corresponding to this age value are 18 - 45 years old and 46 - 69 years old.
[0081] Correspondingly, the confidence type can be determined through the confidence interval. If the age value in the age information is within the preset edge sub-interval, it can be determined that this age information is within the low confidence interval, that is, the confidence type of the age information is the low confidence type. If the age value in the age information is within the preset central sub-interval, it can be determined that this age information is within the high confidence interval, that is, the confidence type of the age information is the high confidence type.
[0082] This alternative embodiment improves the determination mechanism of the target service mode by determining the confidence type and determining the target service mode according to the confidence type, improves the accuracy of the determination result of the target age interval, thereby improving the matching degree between the target service mode and the initiator of the voice information, and further improving the user experience.
[0083] In an alternative embodiment, if the confidence type is a high-confidence type, the age range to which the age information belongs is selected from the adjacent age ranges as the target age range; and the service mode corresponding to the target age range is used as the target service mode.
[0084] Exemplarily, if the confidence type is a high-confidence type, it indicates that the accuracy of determining the age information is relatively high, or the controversy of directly determining the service mode based on the age information is relatively low. Therefore, the age range to which the age information in the adjacent age ranges belongs can be used as the target age range, and then the service mode corresponding to the target age range is used as the target service mode. For example, if the age value in the age information is 8 years old, which belongs to the high-confidence range of 0-9 years old in the age range of 0-12 years old corresponding to the children's mode, then it is determined that the confidence type of the age information is a high-confidence type, and the age range to which the age information belongs, which is 0-12 years old, is used as the target age range.
[0085] In this alternative embodiment, by determining the age range to which the age information belongs according to the adjacent age ranges, when the confidence type is a high-confidence type, the age range to which the age information belongs is directly used as the target age range, improving the accuracy of the determination result of the target age range and helping to improve the matching degree between the target service mode and the initiator of the voice information.
[0086] In another alternative embodiment, if the confidence type is a low-confidence type, the adjacent age ranges are fed back to the initiator of the voice information; the age range selected by the initiator from the adjacent age ranges is used as the target age range; and the service mode corresponding to the target age range is used as the target service mode.
[0087] If the confidence type is a low-confidence type, it indicates that the accuracy of determining the age information is relatively low, or the controversy of directly determining the service mode based on the age information is relatively high. Therefore, the adjacent age ranges can be fed back to the initiator of the voice information. Specifically, the two adjacent age ranges (the age range to which it belongs and the adjacent age range) associated with the age information corresponding to the low confidence can be fed back to the initiator of the voice information. The initiator of the voice information can select according to actual needs from the received adjacent age ranges. The age range selected by the initiator of the voice information is used as the target age range, and the service mode corresponding to the target age range is used as the target service mode.
[0088] For example, if the age value in the age information is 47 years old, which belongs to the low-confidence interval of 46-49 years old in the middle-aged corresponding age interval of 46-69 years old, then determine that the confidence type of this age information is the low-confidence type, and send the adjacent age intervals including the affiliated age interval of 46-69 years old associated with this age information and the adjacent age interval of 18-45 years old to the initiator of the voice information, and the initiator of the voice information selects the age interval according to actual needs.
[0089] In this alternative embodiment, by feeding back the adjacent age intervals to the initiator of the voice information, and taking the age interval selected by the initiator from the adjacent age intervals as the target age interval, it realizes that the initiator can obtain the target age interval according to his own will when the confidence type is the low-confidence type, improves the flexibility and accuracy of determining the target age interval, and thus helps to improve the matching degree between the determined target service mode and the initiator.
[0090] In another alternative embodiment, if the confidence type is the low-confidence type, then feedback the service mode corresponding to the adjacent age intervals to the initiator of the voice information; take the service mode selected by the initiator as the target service mode.
[0091] If the confidence type is the low-confidence type, it indicates that the accuracy of determining this age information is relatively low, or the controversy of directly determining the service mode based on this age information is relatively high. Therefore, the service mode corresponding to the adjacent age intervals can be fed back to the initiator of the voice information. Specifically, it can be to feed back the service mode corresponding to the affiliated age interval associated with the age information corresponding to the low confidence and the service mode corresponding to the adjacent age interval to the initiator of the voice information together. The initiator of the voice information can select from the received service modes according to actual needs. Take the service mode selected by the initiator of the voice information as the target service mode.
[0092] For example, if the age value in the age information is 47 years old, which belongs to the low-confidence interval of 46-49 years old in the middle-aged corresponding age interval of 46-69 years old, then determine that the confidence type of this age information is the low-confidence type, and send the middle-aged mode corresponding to the affiliated age interval of 46-69 years old associated with this age information, and the youth mode corresponding to the adjacent age interval of 18-45 years old to the initiator of the voice information, and the initiator of the voice information selects the service mode according to actual needs.
[0093] The above technical solution feeds back the service mode to the initiator of the voice information, and the initiator selects the target service mode from the fed-back service modes, realizing that the initiator directly selects the target service mode according to its own will when the confidence type is the low-confidence type, improving the flexibility and accuracy of determining the target service mode, and thus helping to improve the matching degree between the determined target service mode and the initiator.
[0094] Optionally, the target age range can be determined according to the historical behavior data of the initiator of the voice information. Among them, the historical behavior data can include the interval selection frequencies for adjacent age ranges when the confidence type is the low-confidence type. Specifically, if the confidence type is the low-confidence type, the historical behavior data of the current initiator of the voice information can be obtained, and according to the interval selection frequencies of the initiator of the voice information for adjacent age ranges in the historical behavior data, the age range with a larger interval selection frequency is used as the target age range. This optional solution does not require intervening in the user's selection operation when the confidence type is the low-confidence type, but automatically makes a selection through the user's historical behavior data, making the voice interaction process more concise and improving the user experience.
[0095] S250. Determine the target resource according to the resource set associated with the target service mode for output.
[0096] The embodiments of the present disclosure determine the age information through audio features; determine the target service mode according to the age information, realizing the automatic selection operation of the target service mode, thereby improving the convenience of the voice interaction process. At the same time, the above technical solution introduces age information to determine the target service mode, and then determines the target resource according to the resource set associated with the target service mode, enabling the determined target resource to adapt to the age situation of the initiator corresponding to the voice information, thereby improving the matching degree between the target service mode and the initiator of the voice information at the age level, and further enhancing the user experience.
[0097] Based on the above technical solutions, the present disclosure also provides an optional embodiment. In this optional embodiment, the voice interaction method is appended. For parts not detailed in this embodiment, reference can be made to the descriptions of the foregoing embodiments, which will not be elaborated here.
[0098] See Figure 3 A voice interaction method, including:
[0099] S310. Obtain voice information.
[0100] S320. Determine the audio features according to the voice information.
[0101] S330. Determine the additional features according to the voice information.
[0102] Among them, the additional features are used as the basis for determining the target resource, and may include at least one of text content, gender information, etc.
[0103] If the additional information includes text content, the text content in the voice information can be extracted based on voice recognition technology. For example, the voice information can be recognized through a voice recognition platform or voice recognition software to obtain the text content; the voice information can also be input into a pre-trained voice recognition model, and according to the model output result, the text content is determined. Among them, the voice recognition model can be trained by a neural network model pre-constructed based on a large amount of text voice information and corresponding text content labels, and the present disclosure does not make any limitation on the specific network structure of the voice recognition model.
[0104] It should be noted that the text content can be all the data after directly converting the voice information into text data, or at least one keyword obtained by extracting keywords from the conversion result after converting the voice information into text form content, so as to reduce the data volume of the text content.
[0105] If the additional information includes gender information, the audio features can be determined according to the voice information, and the gender information can be determined according to the time domain features and / or frequency domain features in the audio features. Among them, the gender information can include male and female. Specifically, the audio features can be input into a pre-trained gender classification model, and according to the model output result, the gender information is determined. Among them, the gender classification model can be trained by a machine learning model or a deep learning model pre-constructed based on a large number of sample audio features and gender information labels. The present disclosure does not make any limitation on the specific network structure of the gender classification model.
[0106] It should be noted that different additional features can be determined according to the voice information simultaneously or successively, and the present disclosure does not make any limitation on the sequence of the determination processes of the additional features. If it is necessary to determine the age information according to the audio features in advance when determining the target service mode, the present disclosure also does not make any limitation on the sequence of the determination processes of the age information and the additional features. For example, the age information and the gender information of the additional features can be determined successively or simultaneously.
[0107] S340. Determine the target service mode according to the audio features.
[0108] Among them, S330 can be executed before or after S340, or can be executed synchronously or alternately with S340, and the present disclosure does not make any limitation on the execution sequence of S330 and S340.
[0109] S350. Select a target resource from the resource set associated with the target service mode according to the additional features for output.
[0110] Optionally, the target resource can be selected from the resource set associated with the target service mode according to the text content in the additional features, so as to improve the matching degree of the target resource and the voice information initiator at the content level. Specifically, resource matching can be performed on the resource set associated with the target service mode according to the text content, and the resource data with a higher matching degree can be used as the target resource, and the target resource can be fed back to the voice information initiator. Among them, the resource matching process can be implemented by means such as resource label similarity matching, and of course other means can also be used for implementation, and the present disclosure does not make any limitation on this.
[0111] Exemplarily, it can be to calculate the relevance between the keywords of the text content and each resource data in the resource set, and use the resource data with the highest relevance as the target resource. Among them, the keywords of the text content can be automatically extracted through natural language processing technology.
[0112] Optionally, the target resource that matches the gender information in the additional features can be selected from the resource set associated with the target service mode, so as to improve the matching degree of the target resource and the voice information initiator at the gender level. For example, in the hospital physical examination scenario, the physical examination instruction information can be output according to the gender of the examinee.
[0113] It should be noted that the present disclosure does not make any limitation on the output method of the target resource. For example, the output method can be determined according to the resource type of the target resource, and then the target resource can be output to the voice information initiator according to the determined output method. For example, the target resource is displayed through at least one of audio playback, video playback, and interface display.
[0114] In the embodiment of the present disclosure, by determining the additional features according to the voice information and selecting the target resource from the resource set associated with the target service mode according to the additional features, the matching degree of the selected target resource and the voice information initiator is improved, and the user experience is improved.
[0115] Based on the above technical solutions, the present disclosure also provides a preferred embodiment. Refer to Figure 4 A voice interaction method shown in
[0116] S401. In response to a mode configuration request, pre-configure the corresponding relationship between different age groups and service modes through the mode configuration module in the skill platform.
[0117] It should be noted that the above corresponding relationship can be added, deleted, or modified according to actual needs.
[0118] S402. The multi-label grouping module in the skill platform is used to group each resource label in the resource platform and establish a corresponding relationship between the label grouping and the service mode.
[0119] S403. The resource platform sends the resource label update status of local resources to the skill platform.
[0120] Among them, the resource platform can send the resource label update status of local resources to the skill platform in real time or at regular intervals. Among them, the update includes addition, deletion, modification, etc.
[0121] S404. The skill platform updates the resource label grouping according to the update status of the resource labels of the existing resources in the resource platform.
[0122] Among them, the resource label grouping can be updated in real time or at regular intervals, and the present disclosure does not make any limitation thereto.
[0123] S405. The operator of the intelligent voice device initiates voice information;
[0124] S406. The intelligent voice device sends the voice information to the cloud server.
[0125] S407. The cloud server determines the audio feature and the text content according to the voice information.
[0126] S408. The cloud server sends the parsed audio feature and text content to the skill platform.
[0127] S409. The skill platform analyzes the audio feature through the feature identification module to determine the age value.
[0128] S410. The skill platform determines the confidence type of the age value through the dialogue management module according to the age value.
[0129] S411. The skill platform judges whether the confidence type is a high confidence type through the dialogue management module; if so, execute S412A; otherwise, execute S412B.
[0130] S412A. The skill platform takes the age range to which the age value belongs as the target age range through the dialogue management module. Continue to execute S414.
[0131] S412B. The skill platform feeds back the age range to which the age value belongs and the adjacent age ranges of the age value to the intelligent voice device through the dialogue management module; or, feeds back the service mode of the age range to which the age value belongs and the service modes of the adjacent age ranges of the age value to the intelligent voice device; continue to execute S413.
[0132] S413. The intelligent voice device sends the selection result to the skill platform in response to the selection operation. Continue to execute S414.
[0133] S414. The skill platform, through the mode configuration module, takes the service mode corresponding to the target age group as the target service mode according to the correspondence between different age groups and service modes, or through the dialogue management module, takes the service mode corresponding to the selection result as the target service mode.
[0134] S415. The skill platform determines the resource tag grouping corresponding to the target service mode through the tag grouping module to determine the resource set corresponding to the target service mode in the resource platform.
[0135] S416. The skill platform determines the target resource from the resource set corresponding to the target service mode according to the text content.
[0136] S417. The skill platform feeds back the target resource to the intelligent voice device.
[0137] S418. The intelligent voice device outputs the target resource.
[0138] For example, a child sends a voice message of "Play XXX" to the intelligent voice device, where "XXX" is a pornographic movie. The intelligent voice device sends the voice message of "Play XXX" to the cloud server; the cloud server extracts the audio features and text features in the voice message and sends the extraction results to the skill platform; the skill platform, based on the feature recognition module, recognizes that the age corresponding to the audio features is 8 years old, so it determines that 8 years old belongs to the high-confidence type in the children's mode. Therefore, from the resource set corresponding to the children's mode, it searches for the content resource that matches "XXX" as the target resource and outputs it through the intelligent voice device, thus avoiding the playback of resource data inappropriate for children. At the same time, without the child manually or carrying the mode category in the voice message, the automated selection of the mode appropriate to the identity of the initiator can be performed, improving the convenience of the voice interaction process.
[0139] As an implementation of the above voice interaction methods, the present disclosure also provides an optional embodiment of an execution device for implementing each voice interaction method. The execution device can be implemented by software and / or hardware and is specifically configured in an electronic device.
[0140] See further Figure 5 , the voice interaction device 500 includes: a voice information acquisition module 501, an audio feature determination module 502, a target service mode determination module 503, and a target resource determination module 504. Among them,
[0141] The voice information acquisition module 501 is used to acquire voice information.
[0142] The audio feature determination module 502 is used to determine audio features according to the voice information.
[0143] A target service mode determination module 503, configured to determine a target service mode according to the audio feature;
[0144] A target resource determination module 504, configured to determine a target resource for output according to a resource set associated with the target service mode.
[0145] In the embodiment of the present disclosure, by introducing an audio feature, a target service mode is determined, and the content resources available for output are restricted through the resource set of the target service mode, avoiding the discomfort brought to the initiator of the voice information by the output of the content resources in other service modes. In addition, in the embodiment of the present disclosure, the audio feature is directly determined according to the voice information, and then the target service mode is automatically determined without manually inputting the target service mode, reducing the user operation, thereby improving the operation convenience of the voice interaction process and enhancing the user experience.
[0146] In an alternative embodiment, the target service mode determination module 503 includes:
[0147] An age information determination unit, configured to determine age information according to the audio feature;
[0148] A target service mode determination unit, configured to determine the target service mode according to the age information.
[0149] In an alternative embodiment, the target service mode determination unit includes:
[0150] A confidence level type determination subunit, configured to determine a confidence level type of the age information according to the age information and an adjacent age interval of the age information;
[0151] A target service mode determination subunit, configured to determine the target service mode from service modes corresponding to the adjacent age intervals according to the confidence level type. In an alternative embodiment, the target service mode determination subunit includes:
[0152] A first age interval selection subunit, configured to, if the confidence level type is a high confidence level type, select an age interval to which the age information belongs from the adjacent age intervals as the target age interval;
[0153] A first target service mode determination subunit, configured to use the service mode corresponding to the target age interval as the target service mode.
[0154] In an alternative embodiment, the target service mode determination subunit includes:
[0155] A service mode feedback subunit, configured to, if the confidence level type is a low confidence level type, feedback the service modes corresponding to the adjacent age intervals to the initiator of the voice information;
[0156] The second target service mode determination slave unit is used to take the service mode selected by the initiator as the target service mode.
[0157] In an optional embodiment, the target service mode determination subunit includes:
[0158] The adjacent age range feedback slave unit is used to, if the confidence type is a low confidence type, feedback the adjacent age range to the initiator of the voice information;
[0159] The second age range selection slave unit is used to take the age range selected by the initiator from the adjacent age ranges as the target age range;
[0160] The third target service mode determination slave unit is used to take the service mode corresponding to the target age range as the target service mode.
[0161] In an optional embodiment, the apparatus further includes:
[0162] The additional feature determination module is used to determine additional features according to the voice information;
[0163] Wherein, the target resource determination module includes:
[0164] The target resource selection unit is used to select a target resource from the resource set associated with the target service mode according to the additional features for output.
[0165] In an optional embodiment, the additional features include text content and / or gender information.
[0166] In an optional embodiment, the apparatus further includes:
[0167] The candidate resource selection module is used to, for any service mode, select candidate resources from the original resources according to the associated keywords of the service mode;
[0168] The candidate resource addition module is used to add the candidate resources to the resource set of the service mode.
[0169] In an optional embodiment, the candidate resource selection module includes:
[0170] The first candidate resource determination unit is used to exclude the original resources marked with the taboo keywords of the service mode from the original resources to obtain the candidate resources; and / or,
[0171] The second candidate resource determination unit is used to select the original resources marked with the permission keywords of the service mode as the candidate resources.
[0172] The above voice interaction device can execute the voice interaction method provided by any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing each voice interaction method.
[0173] In the technical solution of the present disclosure, the acquisition, storage, and application of the involved voice information, etc., all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0174] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0175] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0176] As Figure 6 shown, the device 600 includes a computing unit 601, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0177] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0178] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the voice interaction method. For example, in some embodiments, the voice interaction method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the voice interaction method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the voice interaction method by any other suitable means (e.g., by means of firmware).
[0179] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0180] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0181] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0182] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0183] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0184] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server can also be a server of a distributed system or a server combined with blockchain.
[0185] Artificial intelligence is a discipline that studies how to make computers simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), including both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.
[0186] Cloud computing refers to a technical system that enables elastic and scalable shared physical or virtual resource pools to be accessed through a network. The resources can include servers, operating systems, networks, software, applications, and storage devices, etc., and the resources can be deployed and managed in a on-demand and self-service manner. Through cloud computing technology, it is possible to provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.
[0187] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0188] The above specific embodiments do not constitute a limitation to the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A voice interaction method, comprising: Obtaining voice information; Determining audio features according to the voice information; Determining age information according to the audio features; Determining a confidence type of the age information according to the age information and an adjacent age range of the age information; Determining a target service mode from service modes corresponding to the adjacent age range according to the confidence type; Determining target resources for output according to a resource set associated with the target service mode; Determining the confidence type of the age information according to the age information and the adjacent age range of the age information, including: If an age value in the age information is within a preset marginal sub-range, determining that the confidence type of the age information is a low confidence type; If the age value in the age information is within a preset central sub-range, determining that the confidence type of the age information is a high confidence type.
2. The method according to claim 1, wherein, The determining the target service mode from service modes corresponding to the adjacent age range according to the confidence type includes: If the confidence type is a high confidence type, selecting the age range to which the age information belongs from the adjacent age range as a target age range; Taking the service mode corresponding to the target age range as the target service mode.
3. The method according to claim 1, wherein, The determining the target service mode from service modes corresponding to the adjacent age range according to the confidence type includes: If the confidence type is a low confidence type, feeding back the service mode corresponding to the adjacent age range to the initiator of the voice information; Taking the service mode selected by the initiator as the target service mode.
4. The method according to claim 1, wherein The determining the target service mode from service modes corresponding to the adjacent age range according to the confidence type includes: If the confidence type is a low confidence type, feeding back the adjacent age range to the initiator of the voice information; Taking the age range selected by the initiator from the adjacent age range as a target age range; Taking the service mode corresponding to the target age range as the target service mode.
5. The method according to any one of claims 1-4, further comprising: Determining additional features according to the voice information; Wherein, the determining target resources for output according to a resource set associated with the target service mode includes: Selecting target resources for output from the resource set associated with the target service mode according to the additional features.
6. The method according to claim 5, wherein The additional features include text content and / or gender information.
7. The method according to claim 1, further comprising: For any service mode, selecting candidate resources from original resources according to associated keywords of the service mode; Adding the candidate resources to the resource set of the service mode.
8. The method according to claim 7, wherein The selecting candidate resources from original resources according to the associated keywords of the service mode includes: Excluding original resources marked with taboo keywords of the service mode from the original resources to obtain the candidate resources; and / or, Selecting original resources marked with permitted keywords of the service mode as the candidate resources.
9. A voice interaction device, comprising: A voice information acquisition module for acquiring voice information; An audio feature determination module for determining audio features according to the voice information; A target service mode determination module for determining a target service mode according to the audio features; A target resource determination module for determining a target resource for output according to the resource set associated with the target service mode; The target service mode determination module includes: An age information determination unit for determining age information according to the audio features; A target service mode determination unit for determining the target service mode according to the age information; The target service mode determination unit includes: A confidence type determination subunit for determining the confidence type of the age information according to the age information and the adjacent age interval of the age information; A target service mode determination subunit for determining the target service mode from the service modes corresponding to the adjacent age intervals according to the confidence type; Among them, the confidence type determination subunit is specifically used for: If the age value in the age information is within a preset marginal sub-interval, determining that the confidence type of the age information is a low confidence type; if the age value in the age information is within a preset central sub-interval, determining that the confidence type of the age information is a high confidence type.
10. The device according to claim 9, wherein, The target service mode determination subunit includes: A first age interval selection subunit for, if the confidence type is a high confidence type, selecting the age interval to which the age information belongs from the adjacent age intervals as the target age interval; A first target service mode determination subunit for using the service mode corresponding to the target age interval as the target service mode.
11. The device according to claim 9, wherein, The target service mode determination subunit includes: A service mode feedback subunit for, if the confidence type is a low confidence type, feeding back the service modes corresponding to the adjacent age intervals to the initiator of the voice information; A second target service mode determination subunit for using the service mode selected by the initiator as the target service mode.
12. The apparatus according to claim 9, wherein, The target service mode determination subunit includes: An adjacent age interval feedback subunit for, if the confidence type is a low confidence type, feeding back the adjacent age intervals to the initiator of the voice information; A second age interval selection subunit for using the age interval selected by the initiator from the adjacent age intervals as the target age interval; A third target service mode determination subunit for using the service mode corresponding to the target age interval as the target service mode.
13. The device according to any one of claims 9-12 further includes: An additional feature determination module for determining additional features according to the voice information; Among them, the target resource determination module includes: A target resource selection unit for selecting a target resource for output from the resource set associated with the target service mode according to the additional features.
14. The device according to claim 13, wherein, The additional features include text content and / or gender information.
15. The device according to claim 9 further includes: A candidate resource selection module, configured to select candidate resources from original resources according to the associated keywords of any service mode for any service mode; A candidate resource addition module, configured to add the candidate resources to the resource set of the service mode.
16. The device according to claim 15, wherein, The candidate resource selection module includes: A first candidate resource determination unit, configured to remove the original resources marked with the taboo keywords of the service mode from the original resources to obtain the candidate resources; and / or, A second candidate resource determination unit, configured to select the original resources marked with the permission keywords of the service mode as the candidate resources.
17. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the voice interaction method according to any one of claims 1-8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the voice interaction method according to any one of claims 1-8.
19. A computer program product, comprising computer programs / instructions, wherein when the computer programs / instructions are executed by a processor, the steps of the voice interaction method according to claim 1 are implemented.
Citation Information
Patent Citations
Resource search method, device, terminal, server, and computer-readable storage medium
CN109255053A
Augmented reality method and device
CN109903392A
Multimedia file processing method and device, terminal and storage medium
CN110472074A