Voice control method for NAS device, NAS device, and system

The user's voice is converted into text information through voice control method, and the semantic recognition model is used to analyze text intention to generate control parameters, which solves the problem of inconvenience of users to control NAS devices in busy states, and achieves more flexible and accurate device operations, improving user experience.

WO2025148215A1PCT designated stage expired Publication Date: 2025-07-17SHENZHEN GREEN CONNECTION TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/093411
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-05-15
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Users cannot easily control NAS devices in a busy state, resulting in inconvenient control and poor user experience.

Method used

The user's voice information is converted into text information through a speech control method, and the semantics and text intention of the text information are analyzed using a preset semantic recognition model, and target control parameters are generated to control the operation of the NAS device.

Benefits of technology

It improves the control flexibility of NAS devices and the convenience of user control, and enhances the accuracy and user experience of control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024093411_17072025_PF_FP_ABST
    Figure CN2024093411_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A voice control method for an NAS device, a system, an NAS device, and a computer storage medium, relating to the technical field of NAS. The method comprises: converting acquired voice information of a user into text information (101); on the basis of a preset semantic recognition model, analyzing semantic information of the text information (102); on the basis of the semantic information, analyzing text intent information of the user, the text intent information comprising a target desired control operation (103); determining whether the target desired control operation satisfies a preset control condition (104); and, when it is determined that the target desired control operation satisfies the preset control condition, generating a target control parameter on the basis of the text intent information, the target control parameter being used for controlling an NAS device to execute a first target operation matching the target control parameter, and the first target operation comprising a first operation of controlling the NAS device and / or a first feedback operation (105). According to the method, the control flexibility of the NAS device can be improved, and the convenience of controlling the NAS device by the user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Voice control method for NAS device, NAS device and system Technical Field

[0001] The present invention relates to the field of NAS technology, and in particular to a voice control method for NAS equipment, a NAS equipment, and a system. Background Art

[0002] NAS (Network Attached Storage) is a centralized storage device. Users can store the storage resources of their mobile terminal devices on the storage medium of NAS devices, which can free up the storage space of mobile terminal devices and provide convenience for users' production and life.

[0003] However, in actual NAS device applications, users often need to access NAS devices through complex manual operations on a designated access device. For example, in some scenarios, when a user needs to control a NAS device but is too busy to go to the designated access device to manually control the NAS device, the user needs to use external assistance to continue controlling the NAS device. This makes it inconvenient for the user to control the NAS device, affecting the user's actual control needs and actual usage experience. Therefore, it is particularly important to propose a technical solution that improves the control flexibility of NAS devices and thus helps improve the convenience of user control of NAS devices.

[0004] Summary of the Invention

[0005] The present invention provides a voice control method for a NAS device, a NAS device, and a system, which can improve the control flexibility of the NAS device, thereby facilitating greater convenience for users in controlling the NAS device.

[0006] In order to solve the above technical problems, the first aspect of the present invention discloses a voice control method for a NAS device, the method comprising:

[0007] Convert the acquired user's voice information into text information;

[0008] Analyzing the semantic information of the text information according to a preset semantic recognition model;

[0009] Analyzing the user's text intention information according to the semantic information, wherein the text intention information includes a target desired control operation;

[0010] Determine whether the target expected control operation satisfies a preset control condition. When it is determined that the target expected control operation satisfies the preset control condition, generate a target control parameter based on the text intention information. The target control parameter is used to control the NAS device to perform a first target operation matching the target control parameter. The first target operation includes a first control NAS device operation and / or a first feedback operation matching the text intention information.

[0011] A second aspect of the present invention discloses a NAS device, comprising:

[0012] a memory storing executable program code;

[0013] a processor coupled to the memory;

[0014] The processor calls the executable program code stored in the memory to execute the voice control method for the NAS device disclosed in the first aspect of the present invention.

[0015] A third aspect of the present invention discloses a computer storage medium storing computer instructions. When the computer instructions are called, they are used to execute the voice control method for the NAS device disclosed in the first aspect of the present invention.

[0016] A fourth aspect of the present invention discloses a voice control system for a NAS device. The voice control system for the NAS device includes at least a voice control device for the NAS device and a NAS device communicatively connected to the voice control device of the NAS device. The NAS device stores a plurality of files, and the voice control device of the NAS device performs control operations on the files according to the voice control method for the NAS device disclosed in the first aspect of the present invention.

[0017] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0018] In an embodiment of the present invention, the acquired user voice information is converted into text information; the semantic information of the text information is analyzed according to a preset semantic recognition model; based on the semantic information, the user's text intention information is analyzed, and the text intention information includes a target expected control operation; it is determined whether the target expected control operation meets the preset control condition. When it is determined that the target expected control operation meets the preset control condition, a target control parameter is generated according to the text intention information. The target control parameter is used to control the NAS device to perform a first target operation matching the target control parameter. The first target operation includes a first control NAS device operation and / or a first feedback operation matching the text intention information. It can be seen that the implementation of the present invention can analyze the converted text information according to the preset semantic recognition model, obtain the semantic information of the text information, improve the comprehensibility of the text information, and further analyze the user's text intention information according to the semantic information, improve the analysis feasibility and analysis accuracy of the text intention information, and is conducive to further combining the text intention information to generate target control parameters, improve the control flexibility of the NAS device, and thus help improve the convenience of the user to control the NAS device; at the same time, before generating the target control parameters according to the text intention information, it is judged whether the target expected control operation in the text intention information meets the preset control conditions. When it is judged that the target expected control operation meets the preset control conditions, the operation of generating the target control parameters according to the text intention information is triggered. This can further improve the accuracy of the user's control of the NAS device on the basis of ensuring a full understanding of the text intention information, which is conducive to a better NAS device usage experience for the user. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0020] FIG1 is a flow chart of a method for voice control of a NAS device disclosed in an embodiment of the present invention;

[0021] FIG2 is a flow chart of another method for voice control of a NAS device disclosed in an embodiment of the present invention;

[0022] FIG3 is a schematic structural diagram of a NAS device disclosed in an embodiment of the present invention;

[0023] FIG4 is a schematic diagram of the structure of a voice control system of a NAS device disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] The present invention discloses a voice control method for NAS devices, a NAS device, and a system. The method can analyze converted text information according to a preset semantic recognition model to obtain semantic information of the text information, thereby improving the comprehensibility of the text information, and further analyzing the user's text intention information based on the semantic information, thereby improving the feasibility and accuracy of the analysis of the text intention information, facilitating further combining the text intention information to generate target control parameters, improving the control flexibility of the NAS device, and thereby facilitating improving the convenience of the user controlling the NAS device. At the same time, before generating the target control parameters based on the text intention information, it is determined whether the target expected control operation in the text intention information meets the preset control conditions. When it is determined that the target expected control operation meets the preset control conditions, the operation of generating the target control parameters based on the text intention information is triggered. This can further improve the accuracy of the user controlling the NAS device on the basis of ensuring that the text intention information is fully understood, thereby facilitating a better user experience of the NAS device. The following are detailed descriptions.

[0025] Example 1

[0026] Please refer to Figure 1, which is a flow chart of a method for voice control of a NAS device disclosed in an embodiment of the present invention. The method for voice control of a NAS device described in Figure 1 can be applied to a NAS device, or to an intelligent device related to the NAS device, such as an intelligent device that controls the NAS device, which includes but is not limited to one or more of a cloud device, an edge computing device, a relay device, a base station device, an urban management device, and an intelligent network device. This is not limited in the embodiment of the present invention. As shown in Figure 1, the method for voice control of a NAS device may include the following operations:

[0027] 101. Convert the acquired user's voice information into text information.

[0028] In the embodiment of the present invention, the text information may be specific characters corresponding to the voice information, or may be structured characters that can be understood by a computer and converted based on the voice information.

[0029] Optionally, the above-mentioned conversion tool can be a basic model or an integrated tool, such as FunASR, ENRIE, GPT, etc.

[0030] 102. Analyze the semantic information of the text information according to the preset semantic recognition model.

[0031] 103. Analyze the user's textual intention information based on semantic information, where the textual intention information includes target expected control operations.

[0032] In an embodiment of the present invention, as an optional implementation manner, the above text information includes at least one text sub - information, and analyzing the semantic information of the text information according to the preset semantic recognition model may include the following operations:

[0033] For each text sub - information, analyze the semantic sub - information of the text sub - information according to the preset semantic recognition model.

[0034] And, optionally, analyzing the text intention information of the user according to the semantic information may include the following operations:

[0035] For each text sub - information, determine the sub - information attribute of the text sub - information, and the sub - information attribute is used to represent that the text sub - information is at least one of a noun, a verb, an adjective, an adverb, a preposition, a conjunction, a pronoun, a numeral, an interjection, a sound word, an abbreviation, and an idiom.

[0036] According to the sub - information attribute and the semantic sub - information, analyze the text intention sub - information of the text sub - information.

[0037] According to all the text intention sub - information, analyze the text intention information of the user.

[0038] It can be seen that implementing this optional embodiment can screen the inner - word - property attributes of each text sub - information to improve the analysis accuracy of the text intention sub - information of the text sub - information, and further facilitate improving the analysis accuracy of the text intention information. On the basis of improving the control flexibility of the NAS device, it can further improve the control accuracy of the NAS device and is conducive to improving the user's actual experience of using the NAS device.

[0039] In this optional embodiment, for example, based on the voice information sent by the user, the first - converted text information is: Please play the movie Django Unchained for me. Among them, optionally, each character can represent one of the above text sub - information, or a word, or a noun, such as "Please", "Please play for me", "Play", "Play the movie for me", "Play the movie", "The movie is", etc.

[0040] For the case where the above - mentioned text sub - information is "The movie is", after analyzing the semantic sub - information of the text sub - information for each text sub - information according to the preset semantic recognition model, the following operations may further be included:

[0041] According to the semantic sub - information, judge whether there is a target attribute that does not match in the sub - information attribute of the text sub - information. When it is judged that there is a target attribute in the sub - information attribute, screen out the target text corresponding to the target attribute in the text sub - information and update the text sub - information.

[0042] Furthermore, optionally, the above-mentioned target text is determined as new independent text sub-information.

[0043] It can be seen that the implementation of this optional embodiment can further solve the problem of low accuracy of text intent sub-information analysis caused by mismatch of sub-information attributes in text sub-information for each text sub-information, thereby improving the teaming efficiency of text sub-information and improving the semantic sub-information analysis efficiency and text intent sub-information analysis efficiency of text sub-information, and further ensuring the analysis accuracy of text intent sub-information, which is conducive to improving the analysis accuracy of text intent information, and is conducive to further improving the control accuracy of NAS devices on the basis of improving the control flexibility of NAS devices, and is conducive to improving the user's actual NAS device usage experience.

[0044] In this optional embodiment, as another optional implementation manner, the above-mentioned analysis of the user's text intent information based on all text intent sub-information may include the following operations:

[0045] According to each sub-information attribute, the priority value of each text sub-information is determined.

[0046] Among all the text sub-information, at least one first target text sub-information having a priority value greater than or equal to a preset priority threshold is determined.

[0047] The user's text intention information is analyzed according to the text intention sub-information of all first target text sub-information and the word order of the voice information.

[0048] In this optional embodiment, for example, based on the voice message sent by the user, the text message obtained by conversion is: Please help me mute. In this text message, if each word is a text sub-information, then "quiet" is an adjective. Because it is combined with "sound" to form a verb, the joint priority of the "mute" phrase is higher than other words, so the text intention information of the text message is determined to be "mute". If the aforementioned "help me" is added, the accuracy of determining the text intention information can be further improved. That is, when determining at least one first target text sub-information, this optional embodiment will recombine the text sub-information, and the principle of combination is the proximity principle.

[0049] It can be seen that the implementation of this optional embodiment can set a corresponding priority for each sub-information attribute to determine the priority value of each text sub-information, so that at least one text sub-information with a priority value greater than or equal to the preset priority threshold is screened out among all text sub-information, thereby further combining the text intent sub-information of the text sub-information and the word order of the voice information to analyze the user's text intent information, which can further improve the analysis efficiency of the text intent information on the basis of ensuring the accuracy of the text intent information analysis.

[0050] In this optional embodiment, as another optional implementation, the sorting order of the above-mentioned text sub-information matches the word order of the voice information. The above-mentioned analysis of the user's text intent information based on all text intent sub-information may include the following operations:

[0051] According to the word order of the voice information, the first text sub-information in the text information is determined among all the text sub-information, and the first text sub-information is determined as the current text sub-information.

[0052] The user's predicted text intention information is analyzed based on the text intention sub-information of the current text sub-information and the text intention sub-information of the subsequent adjacent text sub-information of the current text sub-information.

[0053] The subsequent adjacent text sub-information is determined as the new current text sub-information.

[0054] According to the predicted text intent information and the text intent sub-information of the subsequent adjacent text sub-information of the current text sub-information, the predicted text intent information is updated, and the operation of determining the subsequent adjacent text sub-information as the new current text sub-information is triggered, and the predicted text intent information is updated according to the predicted text intent information and the text intent sub-information of the subsequent adjacent text sub-information of the current text sub-information until the subsequent adjacent text sub-information does not exist, and the updated predicted text intent information is determined as the predicted text intent information.

[0055] It can be seen that the implementation of this optional embodiment can provide a method of analyzing text sub-information one by one, thereby combining the semantics, context, and tone of all text information to analyze the text information to update the predicted text intent information, thereby improving the analysis accuracy of the text intent information, which is conducive to further improving the control accuracy of the NAS device on the basis of improving the control flexibility of the NAS device, and playing a complementary role.

[0056] In an optional embodiment, before determining the updated predicted text intention information as the predicted text intention information, and after determining until the subsequent adjacent text sub-information does not exist, the method may further include the following operations.

[0057] Determine the previous adjacent text sub-information of the current text sub-information as the new current text sub-information, and update the predicted text intent information based on the predicted text intent information and the text intent sub-information of the previous adjacent text sub-information of the current text sub-information, and trigger the execution of the operation of determining the previous adjacent text sub-information of the current text sub-information as the new current text sub-information, and updating the predicted text intent information based on the predicted text intent information and the text intent sub-information of the previous adjacent text sub-information of the current text sub-information, until the first text sub-information is the new current text sub-information, then trigger the execution of the operation of determining the updated predicted text intent information as the predicted text intent information.

[0058] It can be seen that the implementation of this optional embodiment can analyze all text information according to the word order of the voice information to analyze the text intent information, and then further analyze all text information in reverse order to filter out the most valuable reference information in the text information, thereby improving the analysis accuracy of the text intent information, which is beneficial to further improve the control accuracy of the NAS device on the basis of improving the control flexibility of the NAS device, and is beneficial to improving the user's convenience and user experience of the NAS device.

[0059] 104. Determine whether the target expected control operation meets the preset control conditions.

[0060] In another optional embodiment, the above-mentioned determination of whether the target desired control operation satisfies the preset control condition may include the following operations:

[0061] According to the target expected control operation, the user's expected controlled module is determined.

[0062] It is determined whether the NAS device has an expected controlled module, and when it is determined that the NAS device has an expected controlled module, a controlled condition of the expected controlled module is determined.

[0063] It is determined whether the user meets the controlled condition. When it is determined that the user meets the controlled condition, it is determined that the target expected control operation meets the preset control condition.

[0064] It can be seen that the implementation of this optional embodiment can further determine the user's expected controlled module based on the target expected control operation in the analyzed text intention information, such as "mute", so as to determine whether the NAS device has the expected controlled module, and, when the judgment result is yes, further determine whether the user meets the controlled conditions, so as to improve the accuracy of the controlled restrictions of the NAS device, improve the access control security of the NAS device, and help improve the user's usage experience.

[0065] Furthermore, in another optional embodiment, the above-mentioned determination of whether the user meets the controlled condition may include the following operations:

[0066] Get the user's identity information.

[0067] Determine the user's preset control permissions based on the user's identity information.

[0068] It is determined whether the preset control authority matches the preset controlled authority of the desired controlled module. When it is determined that the preset control authority matches the preset controlled authority, it is determined that the user meets the controlled condition.

[0069] In this optional embodiment, the user identity information may optionally include one or more of user voice information, user image information, and user biometric information, and may specifically include one or more of timbre, iris, fingerprint, lip print, and the like.

[0070] In this optional embodiment, the above-mentioned preset control authority can be optionally adjusted at any time to meet the situation where the user is temporarily authorized.

[0071] It can be seen that the implementation of this optional embodiment can further determine the user's preset control permissions based on the obtained user identity information, so that when it is determined that the preset control permissions match the preset controlled permissions of the expected controlled module, it is determined that the user meets the controlled conditions, thereby improving the feasibility of the NAS device control restriction solution and the access control security of the NAS device, which is conducive to further improving the user's usage experience.

[0072] 105. When it is determined that the target expected control operation meets the preset control conditions, target control parameters are generated based on the text intention information. The target control parameters are used to control the NAS device to perform a first target operation matching the target control parameters. The first target operation includes a first control NAS device operation and / or a first feedback operation matching the text intention information.

[0073] In an embodiment of the present invention, the above-mentioned first operation of controlling the NAS device can be an operation that matches the text intent information, such as the operation of controlling the NAS device and its associated devices to mute corresponding to the text intent information "mute". The above-mentioned associated devices include but are not limited to one or more terminal devices, cloud devices, edge computing devices, relay devices, base station devices, urban management devices, and intelligent network devices.

[0074] The above-mentioned first feedback operation can be feedback information of executing the above-mentioned first control NAS device operation, or it can be feedback information that the user does not meet the control conditions or similar information, or it can be communication information of real-time dialogue with the user to improve the user's actual usage experience and the voice interaction warmth of the NAS device.

[0075] It can be seen that the implementation of the embodiment of the present invention can analyze the converted text information according to the preset semantic recognition model, obtain the semantic information of the text information, improve the comprehensibility of the text information, and further analyze the user's text intention information according to the semantic information, improve the analysis feasibility and analysis accuracy of the text intention information, and be conducive to further combining the text intention information to generate target control parameters, improve the control flexibility of the NAS device, and thus help improve the convenience of the user to control the NAS device; at the same time, before generating the target control parameters according to the text intention information, it is judged whether the target expected control operation in the text intention information meets the preset control conditions. When it is judged that the target expected control operation meets the preset control conditions, the operation of generating the target control parameters according to the text intention information is triggered. This can further improve the accuracy of the user's control of the NAS device on the basis of ensuring a full understanding of the text intention information, which is conducive to a better NAS device usage experience for the user.

[0076] Example 2

[0077] Please refer to Figure 2, which is a flow chart of a method for voice control of a NAS device disclosed in an embodiment of the present invention. The method for voice control of a NAS device described in Figure 2 can be applied to a NAS device, or to an intelligent device related to the NAS device, such as an intelligent device that controls the NAS device, which includes but is not limited to one or more of a cloud device, an edge computing device, a relay device, a base station device, an urban management device, and an intelligent network device. This is not limited in the embodiment of the present invention. As shown in Figure 2, the method for voice control of a NAS device may include the following operations:

[0078] 201. Determine target audio information based on the acquired environmental information, where the target audio information includes user voice information, and the environmental information includes environmental audio information.

[0079] In an embodiment of the present invention, optionally, the above-mentioned environmental information may also include one or more of environmental image information, environmental lighting information, environmental temperature information, environmental layout information, environmental object information, environmental air information, environmental humidity information, and environmental biological distribution information.

[0080] In the embodiment of the present invention, optionally, the target audio information may further include various noise information in the environment. The noise information may be used to represent all information that the user does not want to hear.

[0081] 202. Calculate information coverage of the human voice information in the target audio information, where the information coverage includes a target frequency ratio.

[0082] Optionally, it may also be the peak value, average value corresponding to the human voice information and the peak value, average value corresponding to the target audio information.

[0083] 203. Determine whether the information coverage is greater than or equal to a preset coverage.

[0084] 204. When it is determined that the information coverage is greater than or equal to the preset coverage, the user's voice information is determined based on the human voice information.

[0085] 205. Convert the acquired user's voice information into text information.

[0086] 206. Analyze the semantic information of the text information according to the preset semantic recognition model.

[0087] 207. Analyze the user's textual intention information based on the semantic information, where the textual intention information includes the target expected control operation.

[0088] 208. Determine whether the target expected control operation meets the preset control conditions.

[0089] 209. When it is determined that the target expected control operation meets the preset control conditions, target control parameters are generated based on the text intention information. The target control parameters are used to control the NAS device to perform a first target operation that matches the target control parameters. The first target operation includes a first control NAS device operation and / or a first feedback operation that matches the text intention information.

[0090] In the embodiment of the present invention, for the supplementary description of steps 205 to 209, please refer to the specific description of steps 101 to 105 in the first embodiment, which will not be repeated in the embodiment of the present invention.

[0091] It can be seen that the implementation of the embodiment of the present invention can process the ambient audio information to obtain more authentic user voice information, and the voice information has the analytical value of textual intent information, thereby improving the accuracy of voice information determination. At the same time, it also improves the flexibility, inclusiveness and diversity of NAS device voice control, which is conducive to further improving the user's NAS device control convenience and experience.

[0092] In the embodiment of the present invention, as an optional implementation manner, the method may further include the following operations:

[0093] When it is determined that the information coverage is less than the preset coverage, the environmental atmosphere information is predicted based on the environmental information.

[0094] Predict the user's current emotional information based on the environmental atmosphere information.

[0095] Generate emotional intention information based on current emotional information.

[0096] Based on the emotional intention information, an emotional control parameter is generated. The emotional control parameter is used to control the NAS device to perform a second target operation matching the emotional control parameter. The second target operation includes a second NAS device control operation and / or a second feedback operation matching the emotional intention information.

[0097] In this optional embodiment, the above-mentioned environmental atmosphere information can be predicted based on one or more of the above-mentioned environmental information. Further, optionally, it can be perceived. Specifically, the environmental atmosphere information can be predicted by sending a perception signal that matches the environmental information in the environment and receiving a feedback signal corresponding to the perception signal. This method can also be used to further directly predict the user's current emotional information. For example, when there is image information of the user, the user's current emotional information can be predicted based on the user's behavior, expression, conversation, etc.

[0098] Optionally, the above-mentioned perception signal may include a communication perception fusion frame structure.

[0099] Optionally, the above-mentioned determination of the target audio information based on the acquired environmental information may include at least one trigger execution condition to protect the user's privacy.

[0100] Optionally, the trigger conditions may include but are not limited to one or more of keywords, instructions, expressions, gestures, body movements, etc.

[0101] As can be seen, implementing this optional embodiment can further analyze the user's emotional intention by sensing the user's ambient environment and current mood when it is determined that human voice information is not valuable, that is, the information coverage is less than the preset coverage. This allows for tentative control interaction with the user to predict the user's control intention and improve the user's NAS control experience. This also helps the user to obtain higher-quality voice signals in conjunction with the NAS device, thereby more accurately controlling the NAS device.

[0102] Example 3

[0103] Please refer to Figure 3, which is a schematic diagram of the structure of a NAS device disclosed in an embodiment of the present invention. As shown in Figure 3, the NAS device may include:

[0104] The memory 401 stores executable program codes.

[0105] A processor 402 is coupled to the memory 401 .

[0106] The processor 402 calls the executable program code stored in the memory 401 to execute the steps of the voice control method for the NAS device described in the first embodiment or the second embodiment of the present invention.

[0107] Example 4

[0108] An embodiment of the present invention discloses a computer storage medium storing computer instructions. When the computer instructions are called, they are used to execute the steps of the voice control method for a NAS device described in the first embodiment or the second embodiment of the present invention.

[0109] Example 5

[0110] [Corrected 07.06.2024 according to Rule 91] Please refer to Figure 4, which shows a voice control system for a NAS device disclosed in an embodiment of the present invention. As shown in Figure 4, the voice control system of the NAS device includes at least a voice control device of the NAS device and a NAS device that is communicatively connected to the voice control device of the NAS device, and the NAS device stores a plurality of files. The voice control device of the NAS device performs control operations on the files according to the voice control method for the NAS device described in Example 1 or Example 2 of the present invention.

[0111] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, or of course, by means of hardware. Based on this understanding, the above technical solution, in essence, or the portion that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

Claims

1. A voice control method for a NAS device, characterized in that, The method includes: Converting the obtained voice information of the user into text information; Analyzing the semantic information of the text information according to a preset semantic recognition model; Analyzing the text intention information of the user according to the semantic information, where the text intention information includes a target expected control operation; Judging whether the target expected control operation meets a preset control condition. When it is judged that the target expected control operation meets the preset control condition, a target control parameter is generated according to the text intention information. The target control parameter is used to control the NAS device to execute a first target operation matching the target control parameter, and the first target operation includes a first control NAS device operation and / or a first feedback operation matching the text intention information.

2. The voice control method of the NAS device according to claim 1, wherein, The text information includes at least one text sub-information. Analyzing the semantic information of the text information according to a preset semantic recognition model includes: For each text sub-information, analyzing the semantic sub-information of the text sub-information according to a preset semantic recognition model; And, analyzing the text intention information of the user according to the semantic information includes: For each text sub-information, determining the sub-information attribute of the text sub-information, where the sub-information attribute is used to indicate that the text sub-information is at least one of a noun, a verb, an adjective, an adverb, a preposition, a conjunction, a pronoun, a numeral, an interjection, an onomatopoeia, an abbreviation, and an idiom; Analyzing the text intention sub-information of the text sub-information according to the sub-information attribute and the semantic sub-information; Analyzing the text intention information of the user according to all the text intention sub-information.

3. The voice control method of the NAS device according to claim 2, wherein Analyzing the text intention information of the user according to all the text intention sub-information includes: Determining a priority value for each text sub-information according to each sub-information attribute; Determining at least one first target text sub-information whose priority value is greater than or equal to a preset priority threshold among all the text sub-information; Analyzing the text intention information of the user according to the text intention sub-information of all the first target text sub-information and the word order of the voice information.

4. The voice control method of the NAS device according to claim 2, characterized in that, The sorting order of the text sub-information matches the word order of the voice information. Analyzing the text intention information of the user according to all the text intention sub-information includes: Determining the first text sub-information in the text information according to the word order of the voice information, and determining the first text sub-information as the current text sub-information; Analyzing the predicted text intention information of the user according to the text intention sub-information of the current text sub-information and the text intention sub-information of the adjacent subsequent text sub-information of the current text sub-information; Determining the adjacent subsequent text sub-information as the new current text sub-information; Update the predicted text intention information according to the text intention sub-information of the subsequent adjacent text sub-information of the predicted text intention information and the current text sub-information, and trigger the execution of the operation of determining the subsequent adjacent text sub-information as the new current text sub-information. Update the predicted text intention information according to the text intention sub-information of the subsequent adjacent text sub-information of the predicted text intention information and the current text sub-information until the subsequent adjacent text sub-information does not exist, and then determine the updated predicted text intention information as the predicted text intention information; And, before determining the updated predicted text intention information as the predicted text intention information, and after until the subsequent adjacent text sub-information does not exist, the method further includes: Determine the previous adjacent text sub-information of the current text sub-information as the new current text sub-information, and update the predicted text intention information according to the text intention sub-information of the previous adjacent text sub-information of the predicted text intention information and the current text sub-information, and trigger the execution of the operation of determining the previous adjacent text sub-information of the current text sub-information as the new current text sub-information and updating the predicted text intention information according to the text intention sub-information of the previous adjacent text sub-information of the predicted text intention information and the current text sub-information until the first text sub-information is the new current text sub-information, and then trigger the execution of the operation of determining the updated predicted text intention information as the predicted text intention information; 5. The voice control method of the NAS device according to claim 1, wherein The judgment of whether the target expected control operation meets the preset control conditions includes: Determine the expected controlled module of the user according to the target expected control operation; Judge whether the expected controlled module exists in the NAS device. When it is judged that the expected controlled module exists in the NAS device, determine the controlled conditions of the expected controlled module; Judge whether the user meets the controlled conditions. When it is judged that the user meets the controlled conditions, determine that the target expected control operation meets the preset control conditions; And, the judgment of whether the user meets the controlled conditions includes: Obtain the user identity information of the user; Determine the preset control authority of the user according to the user identity information; Judge whether the preset control authority matches the preset controlled authority of the expected controlled module. When it is judged that the preset control authority matches the preset controlled authority, determine that the user meets the controlled conditions.

6. The voice control method of the NAS device according to claim 1, characterized in that Before converting the obtained voice information of the user into text information, the method further includes: Determine target audio information according to the obtained environmental information. The target audio information includes the human voice information of the user, and the environmental information includes environmental audio information; Calculate the information coverage of the human voice information in the target audio information. The information coverage includes the target frequency occupancy ratio; Determine whether the information coverage amount is greater than or equal to a preset coverage amount. When it is determined that the information coverage amount is greater than or equal to the preset coverage amount, determine the user's voice information according to the voice information of the user, and trigger the operation of converting the obtained user's voice information into text information; In addition, the method further includes: When it is determined that the information coverage amount is less than the preset coverage amount, predict the environmental atmosphere information according to the environmental information; Predict the user's current emotional information according to the environmental atmosphere information; Generate emotional intention information according to the current emotional information; Generate an emotional control parameter according to the emotional intention information, where the emotional control parameter is used to control the NAS device to execute a second target operation matching the emotional control parameter, and the second target operation includes a second control NAS device operation and / or a second feedback operation matching the emotional intention information.

7. The voice control method of the NAS device according to claim 2, characterized in that After analyzing the semantic sub-information of each text sub-information according to a preset semantic recognition model, the method further includes: For each text sub-information, determine whether there is a target attribute that does not match in the sub-information attributes of the text sub-information according to the semantic sub-information of the text sub-information. When it is determined that there is the target attribute in the sub-information attributes, screen out the target text corresponding to the target attribute in the text sub-information and update the text sub-information.

8. A NAS device, characterized in that, Include: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory and executes the voice control method of the NAS device according to any one of claims 1-7.

9. A computer storage medium, characterized in that, The computer storage medium stores computer instructions, which are used to execute the voice control method of the NAS device according to any one of claims 1-7 when the computer instructions are called.

10. A voice control system for a NAS device, characterized in that, The voice control system of the NAS device at least includes a voice control device and a NAS device communicatively connected to the voice control device, and a plurality of files are stored on the NAS device. The voice control device executes a control operation on the files according to the voice control method of the NAS device according to any one of claims 1-7.

Citation Information

Patent Citations

  • Wakeup-free voice interaction method and device, equipment and storage medium

    CN109326289A

  • Intention recognition method and device, electronic equipment and storage medium

    CN113642334A

  • Vehicle-mounted voice control system and method

    CN115691492A

  • Equipment control method and device, electronic equipment and storage equipment

    CN116229963A

  • Multi-round inquiry-based intention recognition method and device

    CN116415590A