A method for controlling IoT devices based on large language model
By calculating pronunciation highlighting feature values and using predictive language data models for language correction, the problem of insufficient dialect recognition accuracy in the prior art is solved, and the accuracy and efficiency of equipment control are improved.
Patent Information
- Application Number
- CN202510330060.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The prior art processes the user's instruction input in dialect, and the recognition accuracy is insufficient, resulting in the device being unable to accurately understand the user's instructions, affecting the accuracy and reliability of device control.
By obtaining the target language content, extracting pronunciation language parameters, calculating pronunciation highlighting feature values, marking the device control label, and determining the device control method based on the marking results. If the recognition result is abnormal, language correction is performed based on the predicted language data model, and the number of corrections is adjusted to improve the recognition accuracy.
It improves the accuracy of dialect recognition and the accuracy of device control, solves the problem that the device cannot correctly understand dialect instructions, and improves control efficiency.
Smart Images

Figure CN119851667B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Internet of Things, and in particular to an Internet of Things device control method based on a large language model. Background Art
[0002] In recent years, the field of natural language processing has made significant progress, especially the emergence of large language models represented by BERT and GPT. These models can better understand and generate natural language through large-scale corpus training, providing new possibilities for intelligent interaction of IoT devices. Large language models can understand complex human language instructions, making IoT devices more user-friendly. However, with the rapid development of IoT technology, more and more devices are connected to the network, but the interaction and control between devices still face challenges. Traditional control methods require professional technicians to issue specific instructions, which is complicated to operate and has a poor user experience. Therefore, how to achieve accurate and fast interaction between devices has become an urgent problem to be solved.
[0003] China Patent Publication No.: CN118301099A, discloses a method for controlling Internet of Things messages based on the Actor model, the method comprising: S11 receiving control task request information and adding it to the end of the task queue of the pre-built Actor model; S12 the Actor model creates one or more sending Actors; S13 using the ActorManager module to sequentially assign sending Actors to the control task request information; S14 using the assigned sending Actors to issue control task request information; the sending Actor sends its own working message to other sending Actors, exchanges information with other sending Actors, and updates the corresponding status information in the sending Actor itself; S15 when the sending Actor finishes executing the task, it is determined that there is no task in the task queue of the Actor model, and a destruction operation is executed. Through this application method, the Internet of Things devices can be efficiently controlled.
[0004] Chinese patent publication number: CN118675504A, discloses a voice control method and system for intelligent products based on the Internet of Things, the method includes: jointly optimizing the preset acoustic model and language model to obtain the voice recognition model of the product to be controlled; identifying the product characteristics of the product to be controlled, constructing a multi-microphone array of the product to be controlled, and collecting the target voice of the product to be controlled; evaluating the voice direction of the target voice, constructing the voice beam of the multi-microphone array, performing noise constraints on the target voice, and obtaining the constrained target voice; extracting the acoustic characteristics of the constrained target voice, and using the voice recognition model to convert the constrained target voice into voice text; parsing the voice text to obtain the parsed voice text, identifying the user intention of the voice user, constructing the control instructions of the voice user, and executing the voice control of the product to be controlled. The invention can improve the accuracy of voice control of intelligent products.
[0005] However, there are still the following problems in the prior art:
[0006] In actual situations, when controlling a device, it is often encountered that the user inputs commands in dialect, but the existing technology is insufficient in the recognition accuracy of dialects, resulting in an inability to accurately understand the user's commands, and thus making the device unable to receive clear control commands, affecting the accuracy and reliability of device control. Summary of the invention
[0007] To this end, the present invention provides an Internet of Things device control method based on a large language model, which is used to solve the problem that in actual situations, when controlling a device, a user often inputs commands in a dialect, but the prior art is insufficient in the recognition accuracy of dialects, resulting in an inability to accurately understand the user's commands, and further making the device unable to receive clear control commands, affecting the accuracy and reliability of device control.
[0008] To achieve the above object, the present invention provides an Internet of Things device control method based on a large language model, which includes:
[0009] Step S1, obtaining target language content, extracting pronunciation language parameters, including tone features and syllable features, and calculating pronunciation prominence feature values based on the pronunciation language parameters to mark a device control tag;
[0010] Step S2, in response to the marking result of the device control tag, controlling the device, including determining the target language content corresponding to the pronunciation language parameter to control the device;
[0011] Or, based on the predicted language data model, a predicted pronunciation language parameter is obtained to calculate the predicted pronunciation prominence feature value, and the difference between the pronunciation prominence feature value and the predicted pronunciation prominence feature value is determined to determine the adjustment state of the number of language correction times;
[0012] Step S3, obtaining the adjustment state of the number of language corrections to control the device, including recording the corresponding predicted pronunciation language parameters output by the predicted language data model to control the device;
[0013] Or, obtaining abnormal features in the pronunciation language parameters to determine abnormal feature representation values, determining the number of corrections based on the abnormal feature representation values, and recording the pronunciation language parameters after the prediction language data model is corrected a corresponding number of times to control the device;
[0014] The abnormal features include abnormal tone features and abnormal syllable features.
[0015] Furthermore, the process of extracting the pronunciation language parameters includes:
[0016] Converting the time domain audio signal corresponding to the tone into a frequency domain audio signal;
[0017] Determining each frequency domain peak value of the frequency domain audio signal as the tone feature;
[0018] Determine the fundamental frequency curve of the target language content, and determine the average value and change rate of the fundamental frequency curve;
[0019] Extract the formant frequencies corresponding to the target language content;
[0020] The mean value, rate of change, and formant frequency are used as syllable features.
[0021] Furthermore, the process of calculating the pronunciation prominence feature value based on the pronunciation language parameter includes:
[0022] Determine a difference ratio between the tone feature and a reference tone feature as a tone influence factor;
[0023] Determine the difference ratio between the syllable feature and the reference syllable feature as a syllable influence factor;
[0024] The sum of the tone influence factor and the syllable influence factor is determined as a pronunciation prominence feature value.
[0025] Further, the marking device controls the tag, wherein,
[0026] If the pronunciation highlighting feature value is greater than the pronunciation highlighting feature value threshold, marking the device control tag as an abnormal device control tag;
[0027] If the pronunciation highlighting feature value is less than or equal to the pronunciation highlighting feature value threshold, the device control tag is marked as a non-abnormal device control tag.
[0028] Furthermore, the device control method is determined in response to the marking result of the device control tag, wherein:
[0029] If the tag is a non-abnormal device control tag, determining the target language content corresponding to the pronunciation language parameter to control the device;
[0030] If the label is an abnormal device control label, a predicted pronunciation language parameter is obtained based on a predicted language data model to calculate a predicted pronunciation salient feature value, and the difference between the pronunciation salient feature value and the predicted pronunciation salient feature value is determined to determine the adjustment state of the number of language correction times.
[0031] Furthermore, the difference between the pronunciation prominence feature value and the predicted pronunciation prominence feature value is determined to determine the adjustment state of the number of language corrections, wherein:
[0032] If the difference is greater than a preset difference, it is determined that the number of language corrections needs to be adjusted;
[0033] If the difference is less than or equal to the preset difference, it is determined that the number of language corrections does not need to be adjusted.
[0034] Furthermore, the adjustment state of the language correction times is obtained to control the device, wherein:
[0035] If the number of language corrections does not need to be adjusted, recording the corresponding predicted pronunciation language parameters output by the predicted language data model to control the device;
[0036] If the number of language corrections needs to be adjusted, the abnormal features in the pronunciation language parameters are obtained to determine the abnormal feature representation value, the number of corrections is determined based on the abnormal feature representation value, and the pronunciation language parameters after the predicted language data model is corrected the corresponding number of times are recorded to control the device.
[0037] Furthermore, the process of determining the abnormal characteristic value includes:
[0038] Determining a pitch difference between the abnormal pitch feature and the predicted pitch feature and a syllable difference between the abnormal syllable feature and the predicted syllable feature;
[0039] Determine the ratio of the tone difference value to the reference tone difference value as a tone difference influence factor;
[0040] Determine the ratio of the syllable difference value to the reference syllable difference value as the syllable difference influence factor;
[0041] Determine a weighted sum of the tone difference impact factor and the syllable difference impact factor as an abnormal feature representation value;
[0042] If the difference ratio between the tone feature and the benchmark tone feature is greater than the predetermined difference ratio reference value, the tone feature is determined to be an abnormal tone feature; if the difference ratio between the syllable feature and the benchmark syllable feature is greater than the predetermined difference ratio reference value, the syllable feature is determined to be an abnormal syllable feature.
[0043] Furthermore, the process of determining the number of corrections based on the abnormal characteristic value includes:
[0044] If the abnormal characteristic representation value is greater than the preset abnormal characteristic representation value, increasing the number of corrections;
[0045] If the abnormal characteristic characterization value is less than or equal to the preset abnormal characteristic characterization value, the number of corrections is kept unchanged.
[0046] Furthermore, the abnormal characteristic characterization value is positively correlated with the number of corrections.
[0047] Compared with the prior art, the present invention obtains the target language content, extracts the pronunciation language parameters, calculates the pronunciation highlighting feature value, marks the device control label, determines the target language content corresponding to the pronunciation language parameter, and controls the device; or, obtains the predicted pronunciation language parameter based on the predicted language data model, determines the difference between the pronunciation highlighting feature value and the predicted pronunciation highlighting feature value, and determines the adjustment state of the number of language corrections; records the corresponding predicted pronunciation language parameters output by the predicted language data model; or, obtains the abnormal features in the pronunciation language parameters to determine the abnormal feature characterization value, determines the number of corrections based on the abnormal feature characterization value, and records the pronunciation language parameters after the predicted language data model corrects the corresponding number of times, and controls the device. The accuracy of dialect recognition and the accuracy of device control are improved.
[0048] In particular, the device control labels are marked by calculating the pronunciation salient feature values. In the prior art, large language models usually pre-record some dialects to achieve voice recognition of control instructions. However, due to the diversity and particularity of dialects, the prior art still has limitations when processing certain dialects, resulting in some texts being unable to be accurately recognized, which not only affects the accuracy of voice recognition, but also causes the device to be unable to accurately understand the control instructions, resulting in reduced control efficiency. Based on this, the present invention calculates the pronunciation salient feature values to mark the device control labels, and specifically determines the device control method corresponding to the target language content of different labels to improve the accuracy of dialect recognition and the accuracy of device control.
[0049] In particular, predicted pronunciation language parameters are obtained through a predictive language data model to determine the difference between the predicted pronunciation prominence feature value and the pronunciation prominence feature value, providing a data basis for determining the adjustment state of the number of language corrections. In actual situations, when identifying certain uncommon dialects, a model-based prediction method is usually used. However, for some dialects, it is difficult to ensure the accuracy of the recognition results by only using existing models for prediction. Based on this, the present invention obtains the adjustment state of the number of language corrections to control the device by determining the difference between the predicted pronunciation prominence feature value and the pronunciation prominence feature value, so as to improve the accuracy of dialect recognition and the accuracy of device control.
[0050] In particular, for the target language content that needs to adjust the number of language corrections, the number of corrections is determined by calculating the abnormal feature representation value. In actual situations, if the target language content with dialects is continuously corrected without limiting the number of corrections, it may lead to incomplete correction, and thus the device control instructions cannot be accurately identified. In addition, it may also cause the problem of over-correction, resulting in a large amount of computing power resources wasted. Based on this, the present invention determines the number of corrections based on the abnormal feature representation value, and continuously monitors when correcting the target language content with dialects until the correction is complete, so as to improve the accuracy of dialect recognition and the accuracy of device control. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A schematic diagram of the steps of a method for controlling an Internet of Things device based on a large language model according to an embodiment of the invention;
[0052] Figure 2 It is a logic block diagram of a tag control tag of a marking device according to an embodiment of the invention;
[0053] Figure 3 A logic block diagram of determining a device control method in response to a marking result of a device control tag according to an embodiment of the invention;
[0054] Figure 4 A logic block diagram of obtaining the adjustment state of the number of language corrections to control the device according to an embodiment of the invention;
[0055] Figure 5 This is a logic block diagram for determining the number of corrections based on the abnormal feature representation value according to an embodiment of the invention. DETAILED DESCRIPTION
[0056] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0057] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0058] See also Figure 1 , Figure 1 The method steps of the method for controlling an Internet of Things device based on a large language model according to an embodiment of the invention are shown in FIG. The method for controlling an Internet of Things device based on a large language model according to the present invention comprises:
[0059] Step S1, obtaining the target language content, extracting the pronunciation language parameters, including the tone features and the syllable features, and calculating the pronunciation highlight feature value based on the pronunciation language parameters to mark the device control label. It can be understood that the label is a virtual label and can be marked at any time;
[0060] Step S2, in response to the marking result of the device control tag, controlling the device, including determining the target language content corresponding to the pronunciation language parameter to control the device;
[0061] Or, based on the predicted language data model, a predicted pronunciation language parameter is obtained to calculate the predicted pronunciation prominence feature value, and the difference between the pronunciation prominence feature value and the predicted pronunciation prominence feature value is determined to determine the adjustment state of the number of language correction times;
[0062] Step S3, obtaining the adjustment state of the number of language corrections to control the device, including recording the corresponding predicted pronunciation language parameters output by the predicted language data model to control the device;
[0063] Or, obtaining abnormal features in the pronunciation language parameters to determine abnormal feature representation values, determining the number of corrections based on the abnormal feature representation values, and recording the pronunciation language parameters after the prediction language data model is corrected a corresponding number of times to control the device;
[0064] The abnormal features include abnormal tone features and abnormal syllable features.
[0065] Specifically, there is no limitation on the method of acquiring the target language content, for example, the target language content uttered by the user is acquired by a voice acquisition device, which will not be elaborated here.
[0066] Specifically, there is no limitation on the specific type of the predictive language data model. For example, an existing open source speech-based dialect prediction model can be used to convert dialect speech into Mandarin speech for subsequent further recognition. Of course, those skilled in the art can also make a choice based on actual conditions. This is existing technology and will not be elaborated on.
[0067] It is understandable that the predictive language data model can perform multiple predictions to improve accuracy. In implementation, a single prediction completed by the predictive language data model is used as a correction of the target language content, and the number of corrections is the number of predictions of the predictive language data model.
[0068] Specifically, the process of extracting pronunciation language parameters includes:
[0069] Converting the time domain audio signal corresponding to the tone into a frequency domain audio signal;
[0070] Determining each frequency domain peak value of the frequency domain audio signal as the tone feature;
[0071] Determine the fundamental frequency curve of the target language content, and determine the average value and change rate of the fundamental frequency curve;
[0072] Extract the formant frequencies corresponding to the target language content;
[0073] The mean value, rate of change, and formant frequency are used as syllable features.
[0074] Specifically, the average value, change rate and formant frequency of the fundamental frequency curve can reflect the characteristics of the sound from different angles and determine the characteristics of the syllable;
[0075] The average value of the fundamental frequency curve reflects the overall level of the fundamental frequency of the sound within a period of time. Therefore, in this implementation, the arithmetic mean of all fundamental frequency values on the fundamental frequency curve within the time range corresponding to the target language content is determined as the average value;
[0076] The change rate of the fundamental frequency curve can reflect the dynamic changes in the syllable, so in this implementation, the average value of the ratio of the difference between the fundamental frequency values of several adjacent time points and the time interval within the time range corresponding to the target language content is determined as the change rate;
[0077] The formant frequency is the frequency of vocal tract resonance in a speech signal, and is usually used to describe the timbre characteristics of speech to distinguish syllables. The formant frequency can be determined by any signal analysis method in the prior art, such as cepstrum analysis. Of course, other methods can also be used, which will not be described in detail here.
[0078] Specifically, the process of calculating the pronunciation salient feature value based on the pronunciation language parameter includes:
[0079] Determine a difference ratio between the tone feature and a reference tone feature as a tone influence factor;
[0080] Determine the difference ratio between the syllable feature and the reference syllable feature as a syllable influence factor;
[0081] The sum of the tone influence factor and the syllable influence factor is determined as a pronunciation prominence feature value.
[0082] It can be understood that the pitch feature is the frequency domain peak value corresponding to each frequency domain. Therefore, when calculating the difference ratio between the pitch feature and the reference pitch feature, the difference ratio between the pitch feature and the frequency domain peak value corresponding to each frequency domain in the reference pitch feature is solved, and the mean of the difference ratio is determined as the difference ratio between the pitch feature and the reference pitch feature.
[0083] It can be understood that syllable features include average value, rate of change and resonance peak frequency. Therefore, when calculating the difference ratio between syllable features and benchmark syllable features, the difference ratio between the corresponding average value, rate of change and resonance peak frequency in the tone features and the benchmark tone features is solved respectively, and the mean of the difference ratio is determined as the difference ratio between the syllable features and the benchmark syllable features.
[0084] It can be understood that the difference ratio of two values is the ratio of the absolute difference between the two values to the mean of the two values.
[0085] Specifically, it is understandable that those skilled in the art can collect the dialect language of the corresponding region as corpus according to the application region to calculate the reference tone features and the reference syllable features.
[0086] Specifically, the benchmark pitch feature is calculated in advance, and the historical pitch features of several dialect languages of the same type are obtained in advance to determine the average value of several historical pitch features as the benchmark pitch feature.
[0087] Specifically, the benchmark syllable feature is calculated in advance, and the historical syllable features of several dialect languages of the same type are obtained in advance to determine the average value of several historical syllable features as the benchmark syllable feature.
[0088] See also Figure 2 , Figure 2 The following is a logic block diagram of a marking device control tag of an embodiment of the invention. Specifically, the marking device control tag, wherein:
[0089] If the pronunciation highlighting feature value is greater than the pronunciation highlighting feature value threshold, marking the device control tag as an abnormal device control tag;
[0090] If the pronunciation highlighting feature value is less than or equal to the pronunciation highlighting feature value threshold, the device control tag is marked as a non-abnormal device control tag.
[0091] Specifically, the pronunciation highlighting feature value threshold represents the standard for identifying dialects, so the pronunciation highlighting feature value threshold is set within the interval [0.46, 0.78].
[0092] Specifically, the device control label is marked by calculating the pronunciation salient feature value. In the prior art, a large language model usually pre-records some dialects to realize voice recognition of control instructions. However, due to the diversity and particularity of dialects, the prior art still has limitations when processing certain dialects, resulting in some texts that cannot be accurately recognized. This not only affects the accuracy of voice recognition, but also causes the device to be unable to accurately understand the control instructions, resulting in reduced control efficiency. Based on this, the present invention calculates the pronunciation salient feature value to mark the device control label, and specifically determines the device control method corresponding to the target language content of different labels to improve the accuracy of dialect recognition and the accuracy of device control.
[0093] See also Figure 3 , Figure 3 The present invention is a logic block diagram of determining a device control method in response to a marking result of a device control tag according to an embodiment of the present invention. Specifically, in response to a marking result of a device control tag, a device control method is determined, wherein:
[0094] If the tag is a non-abnormal device control tag, determining the target language content corresponding to the pronunciation language parameter to control the device;
[0095] If the label is an abnormal device control label, a predicted pronunciation language parameter is obtained based on a predicted language data model to calculate a predicted pronunciation salient feature value, and the difference between the pronunciation salient feature value and the predicted pronunciation salient feature value is determined to determine the adjustment state of the number of language correction times.
[0096] Specifically, the difference between the pronunciation prominence feature value and the predicted pronunciation prominence feature value is determined to determine the adjustment state of the number of language corrections, wherein:
[0097] If the difference is greater than a preset difference, it is determined that the number of language corrections needs to be adjusted;
[0098] If the difference is less than or equal to the preset difference, it is determined that the number of language corrections does not need to be adjusted.
[0099] Specifically, the preset difference value represents the accuracy of the predicted language content of the predicted language data model and whether the predicted language content can be directly applied, so the preset difference value is set to be selected within the interval [0.25, 0.45].
[0100] See also Figure 4 , Figure 4 This is a logic block diagram of obtaining the adjustment state of the number of language corrections to control the device according to an embodiment of the invention. Specifically, obtaining the adjustment state of the number of language corrections to control the device, wherein:
[0101] If the number of language corrections does not need to be adjusted, recording the corresponding predicted pronunciation language parameters output by the predicted language data model to control the device;
[0102] If the number of language corrections needs to be adjusted, the abnormal features in the pronunciation language parameters are obtained to determine the abnormal feature representation value, the number of corrections is determined based on the abnormal feature representation value, and the pronunciation language parameters after the predicted language data model is corrected the corresponding number of times are recorded to control the device.
[0103] It is understandable that, in the case where the number of language corrections does not need to be adjusted, the predicted pronunciation language parameters and the target language content need to be stored in the model in correspondence so that the same dialect can be accurately and quickly recognized when encountered later.
[0104] Specifically, predicted pronunciation language parameters are obtained through a predictive language data model to determine the difference between the predicted pronunciation prominence feature value and the pronunciation prominence feature value, providing a data basis for determining the adjustment state of the number of language corrections. In actual situations, when identifying certain uncommon dialects, a model-based prediction method is usually used. However, for some dialects, it is difficult to ensure the accuracy of the recognition results by only using existing models for prediction. Based on this, the present invention obtains the adjustment state of the number of language corrections to control the device by determining the difference between the predicted pronunciation prominence feature value and the pronunciation prominence feature value, so as to improve the accuracy of dialect recognition and the accuracy of device control.
[0105] Specifically, the process of determining the abnormal feature representation value includes:
[0106] Determining a pitch difference between the abnormal pitch feature and the predicted pitch feature and a syllable difference between the abnormal syllable feature and the predicted syllable feature;
[0107] Determine the ratio of the tone difference value to the reference tone difference value as a tone difference influence factor;
[0108] Determine the ratio of the syllable difference value to the reference syllable difference value as the syllable difference influence factor;
[0109] Determine a weighted sum of the tone difference impact factor and the syllable difference impact factor as an abnormal feature representation value;
[0110] If the difference ratio between the tone feature and the benchmark tone feature is greater than the predetermined difference ratio reference value, the tone feature is determined to be an abnormal tone feature; if the difference ratio between the syllable feature and the benchmark syllable feature is greater than the predetermined difference ratio reference value, the syllable feature is determined to be an abnormal syllable feature.
[0111] Specifically, the reference tone difference is calculated in advance, a number of target language contents are acquired in advance to determine a number of historical tone differences, and an average value of the number of historical tone differences is determined as the reference tone difference.
[0112] Specifically, the benchmark syllable difference is calculated in advance, and a number of target language contents are obtained in advance to determine a number of historical syllable differences, and an average value of the number of historical syllable differences is determined as the benchmark syllable difference.
[0113] See also Figure 5 , Figure 5 The following is a logic block diagram of determining the number of corrections based on the abnormal feature characterization value according to an embodiment of the invention. Specifically, the process of determining the number of corrections based on the abnormal feature characterization value includes:
[0114] If the abnormal characteristic representation value is greater than the preset abnormal characteristic representation value, increasing the number of corrections;
[0115] If the abnormal characteristic characterization value is less than or equal to the preset abnormal characteristic characterization value, the number of corrections is kept unchanged.
[0116] Specifically, the abnormal feature representation value represents whether the corrected target language content can be accurately recognized, so the abnormal feature representation value is set to be selected within the interval [0.12, 0.26].
[0117] Specifically, for the target language content that needs to adjust the number of language corrections, the number of corrections is determined by calculating the abnormal feature representation value. In actual situations, if the target language content with dialects is continuously corrected without limiting the number of corrections, it may lead to incomplete correction, and thus the device control instructions cannot be accurately recognized. In addition, it may also cause the problem of over-correction, resulting in a large amount of computing power resources wasted. Based on this, the present invention determines the number of corrections based on the abnormal feature representation value, and continuously monitors when correcting the target language content with dialects until the correction is complete, so as to improve the accuracy of dialect recognition and the accuracy of device control.
[0118] Specifically, the abnormal characteristic characterization value is positively correlated with the number of corrections.
[0119] In some possible implementations,
[0120] If the abnormal feature characterization value is greater than or equal to the second abnormal feature characterization value comparison threshold, the number of corrections is set to 4;
[0121] If the abnormal feature characterization value is greater than the first abnormal feature characterization value comparison threshold and less than the second abnormal feature characterization value comparison threshold, the number of corrections is set to 3;
[0122] If the deformation interference characterization parameter is less than or equal to the first deformation interference characterization parameter comparison threshold, the correction times are set to 2;
[0123] Among them, the second abnormal feature characterization value comparison threshold is 1.46 times the abnormal feature characterization value threshold, and the first abnormal feature characterization value comparison threshold is 1.2 times the abnormal feature characterization value threshold.
[0124] It is understandable that the number of corrections can be selected according to actual conditions, which will not be elaborated here.
[0125] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0126] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for controlling Internet of Things devices based on a large language model, characterized in that: include: Step S1, obtaining target language content, extracting pronunciation language parameters, including tone features and syllable features, and calculating pronunciation prominence feature values based on the pronunciation language parameters to mark a device control tag; Step S2, in response to the marking result of the device control tag, controlling the device, if the marked device control tag is a non-abnormal device control tag, determining the target language content corresponding to the pronunciation language parameter to control the device; If the marked device control tag is an abnormal device control tag, a predicted pronunciation language parameter is obtained based on the predicted language data model to calculate the predicted pronunciation prominence feature value, and the difference between the pronunciation prominence feature value and the predicted pronunciation prominence feature value is determined to determine the adjustment state of the number of language corrections, and step S3 is performed; Step S3, obtaining the adjustment state of the number of language corrections to control the device, including, if the number of language corrections does not need to be adjusted, recording the corresponding predicted pronunciation language parameters output by the predicted language data model to control the device; If the number of language corrections needs to be adjusted, the abnormal features in the pronunciation language parameters are obtained to determine the abnormal feature representation value, the number of corrections is determined based on the abnormal feature representation value, and the pronunciation language parameters after the prediction language data model is corrected the corresponding number of times are recorded to control the device; The abnormal features include abnormal tone features and abnormal syllable features; The marking device controls the label, wherein, If the pronunciation highlighting feature value is greater than the pronunciation highlighting feature value threshold, marking the device control tag as an abnormal device control tag; If the pronunciation highlighting feature value is less than or equal to the pronunciation highlighting feature value threshold, marking the device control tag as a non-abnormal device control tag; The process of extracting pronunciation language parameters includes: Converting the time domain audio signal corresponding to the tone into a frequency domain audio signal; Determining each frequency domain peak value of the frequency domain audio signal as the tone feature; Determine the fundamental frequency curve of the target language content, and determine the average value and change rate of the fundamental frequency curve; Extract the formant frequencies corresponding to the target language content; The mean value, rate of change, and formant frequency are used as syllable features.
2. The method for controlling IoT devices based on a large language model according to claim 1, characterized in that: The process of calculating the pronunciation salient feature value based on the pronunciation language parameter comprises: Determine a difference ratio between the tone feature and a reference tone feature as a tone influence factor; Determine the difference ratio between the syllable feature and the reference syllable feature as a syllable influence factor; The sum of the tone influence factor and the syllable influence factor is determined as a pronunciation prominence feature value.
3. The method for controlling Internet of Things devices based on a large language model according to claim 1, characterized in that: The device control method is determined in response to the marking result of the device control tag, wherein: If the tag is a non-abnormal device control tag, determining the target language content corresponding to the pronunciation language parameter to control the device; If the label is an abnormal device control label, a predicted pronunciation language parameter is obtained based on a predicted language data model to calculate a predicted pronunciation salient feature value, and the difference between the pronunciation salient feature value and the predicted pronunciation salient feature value is determined to determine the adjustment state of the number of language correction times.
4. The method for controlling IoT devices based on a large language model according to claim 1, characterized in that: The difference between the pronunciation prominence feature value and the predicted pronunciation prominence feature value is determined to determine the adjustment state of the number of language corrections, wherein: If the difference is greater than a preset difference, it is determined that the number of language corrections needs to be adjusted; If the difference is less than or equal to the preset difference, it is determined that the number of language corrections does not need to be adjusted.
5. The method for controlling Internet of Things devices based on a large language model according to claim 4, characterized in that: The adjusting state of the number of language corrections is obtained to control the device, wherein: If the number of language corrections does not need to be adjusted, recording the corresponding predicted pronunciation language parameters output by the predicted language data model to control the device; If the number of language corrections needs to be adjusted, the abnormal features in the pronunciation language parameters are obtained to determine the abnormal feature representation value, the number of corrections is determined based on the abnormal feature representation value, and the pronunciation language parameters after the predicted language data model is corrected the corresponding number of times are recorded to control the device.
6. The method for controlling IoT devices based on a large language model according to claim 1, characterized in that: The process of determining the abnormal characteristic value includes: Determining a pitch difference between the abnormal pitch feature and the predicted pitch feature and a syllable difference between the abnormal syllable feature and the predicted syllable feature; Determine the ratio of the tone difference value to the reference tone difference value as a tone difference influence factor; Determine the ratio of the syllable difference value to the reference syllable difference value as the syllable difference influence factor; Determine a weighted sum of the tone difference impact factor and the syllable difference impact factor as an abnormal feature representation value; Among them, if the difference ratio between the tone feature and the benchmark tone feature is greater than the predetermined difference ratio reference value, the tone feature is determined to be an abnormal tone feature; if the difference ratio between the syllable feature and the benchmark syllable feature is greater than the predetermined difference ratio reference value, the syllable feature is determined to be an abnormal syllable feature.
7. The method for controlling Internet of Things devices based on a large language model according to claim 1, characterized in that: The process of determining the number of corrections based on the abnormal characteristic value includes: If the abnormal characteristic representation value is greater than the preset abnormal characteristic representation value, increasing the number of corrections; If the abnormal characteristic characterization value is less than or equal to the preset abnormal characteristic characterization value, the number of corrections is kept unchanged.
8. The method for controlling Internet of Things devices based on a large language model according to claim 1, characterized in that: The abnormal characteristic characterization value is positively correlated with the number of corrections.
Citation Information
Patent Citations
Internet of Things message control method based on Actor model
CN118301099A
Method and system for realizing voice control of intelligent product based on Internet of Things
CN118675504A
AI voice control method and system for intelligent household electrical appliance
CN119028346A