Smart life AI interaction system and method based on language recognition
Through the smart life AI interaction system based on language recognition, positioning and cameras are used to determine the location of the elderly, monitor voice signals, adjust language recognition model parameters, and detect text similarity, which solves the problem of low efficiency in elderly care and realizes more intelligent AI interaction.
Patent Information
- Application Number
- CN202411647069.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-11-18
AI Technical Summary
How to enhance the intelligence of AI interactions in smart life to improve the efficiency of elderly care, especially when the elderly do not have sufficient companionship.
Through the smart life AI interactive system based on language recognition, positioning sensors and cameras are used to determine the area of the target object, monitor its voice signal, and adjust the parameters of the language recognition model according to identity information and physiological state parameters to perform language recognition, detect the similarity of text content, perform corresponding operations, and combine regional and local characteristics to improve understanding ability.
It improves the intelligence of AI interaction, avoids comprehension bias, ensures that AI better understands the needs and emotional changes of the elderly, and improves the efficiency of elderly care.
Smart Images

Figure CN119580740B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of smart life or the field of Internet of Things technology, and specifically to a smart life AI interaction system and method based on language recognition. Background Art
[0002] With the aging population in society, elderly care has become a hot topic. With the rapid development of electronic technology, smart living has gradually become part of every household, and smart elderly care has become a trend in the times. In real life, the number of elderly people with children accompanying them in their care is decreasing. Therefore, the question of how to improve the efficiency of elderly care by enhancing the intelligence of smart AI interactions is an urgent issue. Summary of the Invention
[0003] The embodiments of the present application provide a smart life AI interaction system and method based on language recognition, which can improve the intelligence of smart life AI interaction and thus improve the efficiency of elderly care.
[0004] In a first aspect, an embodiment of the present application provides a language recognition-based smart life AI interaction method, which is applied to a control platform device in a language recognition-based smart life AI interaction system, wherein the control platform device includes a first language recognition model, and the language recognition-based smart life AI interaction system further includes multiple language recognition systems, each language recognition system corresponding to a region and a second language recognition model, and the language recognition system of each region includes at least one positioning sensor, at least one camera, and at least one voice interaction device. The method includes:
[0005] Positioning the target object by using the at least one positioning sensor and the at least one camera in each area to determine the target area where the target object is located;
[0006] monitoring a first speech signal of the target object by the target language recognition system in the target area;
[0007] Acquiring target identity information and target physiological state parameters of the target object;
[0008] adjusting second model parameters of a target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameter to obtain target second model parameters;
[0009] Performing language recognition on the first speech signal using the target second language recognition model and the target second model parameters to obtain first text content;
[0010] performing language recognition on the first speech signal according to the first language recognition model to obtain second text content;
[0011] detecting a first similarity between the first text content and the second text content;
[0012] When the first similarity is greater than a first preset threshold, a first response operation is performed according to the first text content.
[0013] In a second aspect, an embodiment of the present application provides a language recognition-based smart life AI interaction system, which is applied to a control platform device in the language recognition-based smart life AI interaction system. The control platform device includes a first language recognition model. The language recognition-based smart life AI interaction system also includes multiple language recognition systems, each language recognition system corresponding to a region and a second language recognition model. The language recognition system in each region includes at least one positioning sensor, at least one camera, and at least one voice interaction device. The language recognition-based smart life AI interaction system includes:
[0014] a positioning unit, configured to locate a target object by using the at least one positioning sensor and the at least one camera in each area, so as to determine a target area where the target object is located;
[0015] A monitoring unit, configured to monitor a first speech signal of the target object through the target language recognition system in the target area;
[0016] an acquiring unit, configured to acquire target identity information and target physiological state parameters of the target object;
[0017] An adjusting unit, configured to adjust a second model parameter of a target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameter to obtain a target second model parameter;
[0018] a recognition unit configured to perform language recognition on the first speech signal using the target second language recognition model and the target second model parameters to obtain first text content; and perform language recognition on the first speech signal based on the first language recognition model to obtain second text content;
[0019] a detection unit, configured to detect a first similarity between the first text content and the second text content;
[0020] An execution unit is configured to execute a first response operation according to the first text content when the first similarity is greater than a first preset threshold.
[0021] The implementation of the embodiments of this application has the following beneficial effects:
[0022] It can be seen that the language recognition-based smart life AI interaction system and method described in the embodiments of the present application are applied to a control platform device in a language recognition-based smart life AI interaction system. The control platform device includes a first language recognition model, and the language recognition-based smart life AI interaction system also includes multiple language recognition systems, each language recognition system corresponds to an area and a second language recognition model, and the language recognition system of each area includes at least one positioning sensor, at least one camera, and at least one voice interaction device. The method includes: locating the target object through the at least one positioning sensor and the at least one camera in each area to determine the target area where the target object is located; monitoring the first voice signal of the target object through the target language recognition system in the target area; obtaining target identity information and target physiological state parameters of the target object; adjusting the second model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain target second model parameters; performing language recognition on the first voice signal through the target second language recognition model and the target second model parameters to obtain first text content; performing language recognition on the first voice signal according to the first language recognition model to obtain second text content; detecting a first similarity between the first text content and the second text content;When the first similarity is greater than a first preset threshold, a first response operation is performed according to the first text content. First, the target physiological state parameters reflect the actual needs and emotional changes of the target object to a certain extent. The second model parameters of the target second language recognition model of the target language recognition system can be adjusted according to the target identity information and the target physiological state parameters to obtain the target second model parameters. Obtaining model parameters that are deeply related to the actual situation of the target object helps to improve AI understanding capabilities and ensure the AI interaction capabilities and quality of smart life of smart life. Second, since the second language recognition model has regional characteristics, the recognition effect of the first text content also conforms to the regional characteristics. The first voice signal is language recognized by the target second language recognition model and the target second model parameters to obtain the first text content. The first text content has more regional characteristics. Third, since the first language recognition model can be based on the perspective of the target object's interactive object The training process also considers the language system (local characteristics) related to the target object. The first language recognition model is more locally specific and fully considers social (family and friends) understanding, not restricted by a single region. Fourthly, the first similarity between the first text content and the second text content is detected. To a certain extent, the similarity between the two includes the regional characteristics of the model, the local characteristics, and the full consideration of social (family and friends) understanding. When the first similarity is greater than a first preset threshold, it indicates that the model's understanding of the target object's language is deeply consistent with the regional characteristics and local characteristics, and fully considers social (family and friends) understanding. Furthermore, a first response operation can be executed based on the first text content. This not only avoids misunderstandings during AI interaction, but also ensures that AI is more understanding, thereby enhancing the intelligence of the smart life AI interaction system based on language recognition, thereby improving the efficiency of elderly care. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 This is a structural diagram of a smart life AI interaction system based on language recognition provided by an embodiment of the present application;
[0025] Figure 2 This is a flow chart of a method for intelligent life AI interaction based on language recognition provided by an embodiment of the present application;
[0026] Figure 3This is a schematic diagram of the structure of a control platform device provided in an embodiment of the present application;
[0027] Figure 4 This is a block diagram of the functional units of another smart life AI interaction system based on language recognition provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0029] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0030] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0031] The control platform device described in the embodiments of the present application may be a server, such as an edge server or a local server. The control platform device may also include other devices.
[0032] The following is a detailed introduction to the embodiments of the present application.
[0033] See also Figure 1 , Figure 1 This is a structural diagram of a language recognition-based smart life AI interaction system provided in an embodiment of the present application. The language recognition smart life AI interaction system includes: a control platform device, multiple language recognition systems, each language recognition system corresponds to a region and a second language recognition model, and the language recognition system in each region includes at least one positioning sensor, at least one camera, and at least one voice interaction device.
[0034] The control platform device is in communication connection with each of the multiple language recognition systems.
[0035] In specific implementation, the above-mentioned language recognition smart life AI interaction system can be used to achieve the following functions:
[0036] Positioning the target object by using the at least one positioning sensor and the at least one camera in each area to determine the target area where the target object is located;
[0037] monitoring a first speech signal of the target object by the target language recognition system in the target area;
[0038] Acquiring target identity information and target physiological state parameters of the target object;
[0039] adjusting second model parameters of a target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameter to obtain target second model parameters;
[0040] Performing language recognition on the first speech signal using the target second language recognition model and the target second model parameters to obtain first text content;
[0041] performing language recognition on the first speech signal according to the first language recognition model to obtain second text content;
[0042] detecting a first similarity between the first text content and the second text content;
[0043] When the first similarity is greater than a first preset threshold, a first response operation is performed according to the first text content.
[0044] Optionally, the language recognition smart life AI interaction system is specifically used for:
[0045] When the first similarity is less than or equal to the first preset threshold and greater than a second preset threshold, acquiring, by the target language recognition system in the target area, video content of a preset time period corresponding to the first voice signal;
[0046] Performing action recognition on the video content to obtain third text content, wherein the third text content is used to represent the action intention or action state of the target object;
[0047] determining a second similarity between the first text content and the third text content;
[0048] determining a third similarity between the second text content and the third text content;
[0049] determining a larger value between the second similarity and the third similarity, and when the larger value is greater than a set threshold, performing a second response operation according to text content corresponding to the larger value;
[0050] When the larger value is less than or equal to the set threshold, the target language recognition system in the target area prompts the target subject to re-input voice content.
[0051] Optionally, the language recognition smart life AI interaction system is specifically used for:
[0052] When the first similarity is less than or equal to the second preset threshold, the step of prompting the target object to re-input voice content through the target language recognition system in the target area is performed.
[0053] Optionally, in terms of adjusting the model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain the target model parameters, the language recognition smart life AI interaction system is specifically used to:
[0054] Determining second initial model parameters corresponding to the target second language recognition model corresponding to the target identity information;
[0055] determining a first adjustment parameter corresponding to the target physiological state parameter;
[0056] The second initial model parameters are adjusted according to the first adjustment parameters to obtain the target second model parameters.
[0057] Optionally, the language recognition smart life AI interaction system is specifically used for:
[0058] Determining target sample data corresponding to the target area and the target identity information;
[0059] Using the target sample data to train the first language recognition model to obtain a trained first language recognition model;
[0060] The performing language recognition on the first speech signal according to the first language recognition model to obtain second text content includes:
[0061] determining first initial model parameters corresponding to the first speech recognition model corresponding to the target physiological state parameter;
[0062] determining a first signal-to-noise ratio relative to the first speech signal;
[0063] determining a second adjustment parameter corresponding to the first signal-to-noise ratio;
[0064] Adjusting the first initial model parameters according to the second adjustment parameters to obtain the target first model parameters;
[0065] Perform language recognition on the first speech signal according to the trained first language recognition model and the target first model parameters to obtain the second text content.
[0066] It can be seen that the smart life AI interaction system based on language recognition described in the embodiment of the present application is applied to the control platform equipment. First, the target physiological state parameters reflect the actual needs and emotional changes of the target object to a certain extent. The second model parameters of the target second language recognition model of the target language recognition system can be adjusted according to the target identity information and the target physiological state parameters to obtain the target second model parameters. The model parameters that are deeply related to the actual situation of the target object are obtained, which helps to improve the AI comprehension ability and ensure the smart life AI interaction ability and the quality of smart life. Second, since the second language recognition model has regional characteristics, the recognition effect of the first text content also conforms to the regional characteristics. The first voice signal is language recognized by the target second language recognition model and the target second model parameters to obtain the first text content. The first text content has more regional characteristics. Third, since the first language recognition model can be based on the interactive object of the target object , and also consider more about the language system (local characteristics) related to the target object. The first language recognition model is more local and fully considers the understanding of social (relatives and friends), and is not restricted by the constraints of one region. Fourthly, the first similarity between the first text content and the second text content is detected. The similarity between the two is to a certain extent, that is, the regional characteristics of the model included, and the local characteristics and full consideration of social (relatives and friends) understanding. When the first similarity is greater than the first preset threshold, it means that the model's language understanding of the target object is deeply consistent with the regional characteristics and the local characteristics and full consideration of social (relatives and friends) understanding. Then, the first response operation can be performed according to the first text content, thereby not only avoiding the understanding deviation in the AI interaction process, but also ensuring that AI is more considerate, and then, the intelligence of the smart life AI interaction system based on language recognition can be improved to improve the efficiency of elderly care.
[0067] See also Figure 2 , Figure 2This is a flow chart of a language recognition-based smart life AI interaction method provided by an embodiment of the present application. As shown in the figure, a control platform device is applied to a language recognition-based smart life AI interaction system. The control platform device includes a first language recognition model. The language recognition-based smart life AI interaction system also includes multiple language recognition systems. Each language recognition system corresponds to a region and a second language recognition model. The language recognition system of each region includes at least one positioning sensor, at least one camera, and at least one voice interaction device. This language recognition-based smart life AI interaction method includes:
[0068] 201. Position a target object using the at least one positioning sensor and the at least one camera in each area to determine a target area where the target object is located.
[0069] In a specific implementation, the control platform device includes a first language recognition model, which includes first model parameters.
[0070] The first language recognition model and the second language recognition model may each include at least one of the following: a large language model (such as GPT, NLP), a neural network model, Speech Synthesis Markup Language (SSML), etc., which are not limited here.
[0071] The target object may be a person, for example, an elderly person, and the elderly person may be understood as a group of people older than a preset age, and the preset age may be pre-set or a system default.
[0072] The positioning sensor may be a proximity sensor, a laser sensor, an ultrasonic sensor, etc., which are not limited here. The voice interaction device may include at least one sound sensor and at least one microphone.
[0073] Among them, the smart life AI interaction system based on language recognition also includes multiple language recognition systems, each language recognition system corresponds to an area and a second language recognition model, and the language recognition system of each area includes at least one positioning sensor, at least one camera and at least one voice interaction device. The area can be divided based on the indoor map. For example, each independent living area is regarded as an area, for example, the toilet is an area, room 1 is an area, room 2 is an area, the kitchen is an area, and so on. For example, the area is divided based on the user's activity arrangements. For example, the meal time corresponds to an area, and the sleeping time corresponds to an area.
[0074] In a specific implementation, each region corresponds to a second language recognition model, and each second language recognition model corresponds to a second model parameter. The second language recognition model can be trained based on the speech signal collected in the region, so that the second language recognition model has regional characteristics, and different regions correspond to different characteristics.
[0075] In a specific implementation, the target object can be positioned by at least one positioning sensor and at least one camera in each area to determine the target area where the target object is located. The positioning sensor can locate the user's position, and the camera can correspond to the orientation and corresponding action direction.
[0076] Among them, the first language recognition model can be trained based on at least one of the following sample data: voice data of the target object, voice data of other characters who have conversations with the target object, conversation content related to the target object, voice data of other characters corresponding to the target object's social circle, etc., which are not limited here.
[0077] 202. Monitor a first speech signal of the target object through the target language recognition system in the target area.
[0078] In the embodiment of the present application, the first speech signal of the target object is monitored by the target language recognition system in the target area, so that the speech of the target object can be monitored immediately.
[0079] In which, the target object only exists in one area in each time period, and the first voice signal can be collected by only using the voice interaction device in the area where the target object is located, thereby reducing system power consumption.
[0080] 203. Obtain target identity information and target physiological state parameters of the target object.
[0081] In an embodiment of the present application, the target identity information may include at least one of the following: age, gender, place of birth, occupation, resume, hobbies, social circle, family relationship, medical condition, IQ, EQ, hearing level, language expression ability, etc., which are not limited here.
[0082] The target physiological state parameter may include at least one of the following: blood pressure, blood sugar, heart rate, blood temperature, respiratory rate, blood lipids, body temperature, etc., without limitation. The target subject may wear a wearable device, through which the target physiological state parameter may be obtained, and the wearable device may be communicatively connected to the control platform device.
[0083] 203. Adjust second model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain target second model parameters.
[0084] In the embodiment of the present application, different identity information corresponds to different language systems, and its corresponding social circles and cognitive understanding levels are different. Therefore, different models are required for language recognition. The target physiological state parameters reflect the actual needs and emotional changes of the target object to a certain extent. In this way, the second model parameters of the target second language recognition model of the target language recognition system can be adjusted according to the target identity information and the target physiological state parameters to obtain the target second model parameters. The model parameters that are deeply related to the actual situation of the target object are obtained, which helps to improve the AI comprehension ability and ensure the AI interaction ability and quality of smart life.
[0085] 204. Perform language recognition on the first speech signal using the target second language recognition model and the target second model parameters to obtain first text content.
[0086] In the embodiment of the present application, since the second language recognition model has regional characteristics, the recognition effect of the first text content also conforms to the regional characteristics. The first speech signal is subjected to language recognition through the target second language recognition model and the target second model parameters to obtain the first text content, which has more regional characteristics.
[0087] 205. Perform language recognition on the first speech signal according to the first language recognition model to obtain second text content.
[0088] In the embodiment of the present application, since the first language recognition model can be trained based on the perspective of the target object's interactive object and also takes more account of the language system (local characteristics) related to the target object, the first language recognition model is more local and fully considers social (relatives and friends) understanding, and is not restricted by the constraints of a region.
[0089] In a specific implementation, the first speech signal can be subjected to language recognition according to the first language recognition model to obtain the second text content. That is, the first language recognition model is compatible with the understanding of interactive objects related to the target object and local characteristic attributes, so that the second text content is more in line with the target object's life communication scenarios.
[0090] 206. Detect a first similarity between the first text content and the second text content.
[0091] In an embodiment of the present application, the first similarity between the first text content and the second text content can be detected. The similarity between the two, to a certain extent, includes the regional characteristics of the model, the local characteristics, and fully considers the understanding of social (relatives and friends). It can not only avoid understanding deviations in the AI interaction process, but also ensure that AI is more considerate. Furthermore, it can enhance the intelligence of the smart life AI interaction system based on language recognition to improve the efficiency of elderly care.
[0092] 207. When the first similarity is greater than a first preset threshold, perform a first response operation according to the first text content.
[0093] The first preset threshold may be preset or set by system default.
[0094] In a specific implementation, when the first similarity is greater than the first preset threshold, it indicates that the model's understanding of the language of the target object is deeply consistent with regional characteristics and local characteristics, and fully considers social (relatives and friends) understanding. Therefore, the first response operation can be performed according to the first text content.
[0095] Among them, the first response operation can be to execute a certain instruction, such as turning on the light, flushing the toilet, asking for help, etc. The first response operation can also be to reply to the voice content in the tone of a certain character.
[0096] In a specific implementation, when performing the first response operation according to the first text content, different physiological state parameters and text contents require different reply roles. In a specific implementation, the preset physiological state parameters and the mapping relationship between the text content and the reply role can be pre-stored. Then, the target reply role corresponding to the physiological state parameter and the text content can be determined based on the mapping relationship, and the target reply content corresponding to the target reply role and the first text content is generated. The first response operation is performed based on the target reply role with the target reply content. The reply role can include at least one of the following: doctor, caregiver, lawyer, child, spouse, friend, etc., which are not limited here.
[0097] Furthermore, the reply process corresponding to the reply role can also be recorded to generate a corresponding video file, and the video file can be sent to the reply role.
[0098] Optionally, the following steps may also be included:
[0099] When the first similarity is less than or equal to the first preset threshold and greater than a second preset threshold, acquiring, by the target language recognition system in the target area, video content of a preset time period corresponding to the first voice signal;
[0100] Performing action recognition on the video content to obtain third text content, wherein the third text content is used to represent the action intention or action state of the target object;
[0101] determining a second similarity between the first text content and the third text content;
[0102] determining a third similarity between the second text content and the third text content;
[0103] determining a larger value between the second similarity and the third similarity, and when the larger value is greater than a set threshold, performing a second response operation according to text content corresponding to the larger value;
[0104] When the larger value is less than or equal to the set threshold, the target language recognition system in the target area prompts the target subject to re-input voice content.
[0105] In the embodiment of the present application, when the first similarity is less than or equal to a first preset threshold and greater than a second preset threshold, it indicates that the model's understanding of the target object's language generally conforms to regional characteristics and local characteristics and fully considers social (relatives and friends) understanding. The first preset threshold is greater than the second preset threshold.
[0106] The preset time period may be pre-set or set by the system as a default. For example, the preset time period may include a start time and an end time of the first voice signal.
[0107] The threshold value may be preset or set by system default.
[0108] Furthermore, the target language recognition system of the target area can be used to obtain the video content of the preset time period corresponding to the first voice signal, and then the action recognition is performed through the video content to obtain the third text content, which is used to characterize the action intention or action state of the target object. Then, the second similarity between the first text content and the third text content can be determined, and the third similarity between the second text content and the third text content can be determined. The larger value of the second similarity and the third similarity is determined. When the larger value is greater than the set threshold, it indicates that the user intention reflected by the action is consistent with the user intention reflected by the voice. Figure 1If the value is the same as the set threshold, the second response operation can be performed according to the text content corresponding to the larger value. When the larger value is less than or equal to the set threshold, it means that the user intention reflected by the action is inconsistent with the user intention reflected by the voice. The target language recognition system in the target area prompts the target object to re-enter the voice content. That is, when the model's understanding of the target object's language generally conforms to regional characteristics and local characteristics and fully considers social (relatives and friends) understanding, action recognition can be used to deeply analyze the user's intention to determine the user's true thoughts. This can not only avoid comprehension deviations in the AI interaction process, but also ensure that AI is more considerate, and further, it can enhance the intelligence of the smart life AI interaction system based on language recognition to improve the efficiency of elderly care.
[0109] In a specific implementation, when performing the second response operation according to the text content corresponding to the larger value, since different physiological state parameters and text contents require different reply roles, in a specific implementation, the preset physiological state parameters and the mapping relationship between the text content and the reply role can be pre-stored. Then, the target reply role corresponding to the physiological state parameters and the text content can be determined based on the mapping relationship, and the target reply content corresponding to the target reply role and the text content corresponding to the larger value is generated. The second response operation is performed based on the target reply role with the target reply content. The reply role can include at least one of the following: doctor, caregiver, lawyer, child, spouse, friend, etc., which are not limited here.
[0110] In a specific implementation, when the larger value is less than or equal to the set threshold, different physiological state parameters require different response roles. In a specific implementation, the mapping relationship between the preset physiological state parameters and the response roles can be pre-stored, and then the response role corresponding to the physiological state parameter can be determined based on the mapping relationship. The target language recognition system in the target area can use the response role to prompt the target object to re-enter the voice content.
[0111] Furthermore, the reply process corresponding to the reply role can also be recorded to generate a corresponding video file, and the video file can be sent to the reply role.
[0112] Optionally, the following steps may also be included:
[0113] When the first similarity is less than or equal to the second preset threshold, the step of prompting the target object to re-input voice content through the target language recognition system in the target area is performed.
[0114] In a specific implementation, when the first similarity is less than or equal to the second preset threshold, the model's understanding of the target object's language does not conform to regional characteristics and local characteristics, and fully considers social (relatives and friends) understanding. Then, the step of prompting the target object to re-enter the voice content through the target language recognition system in the target area can be executed. This not only avoids understanding deviations during AI interaction, but also ensures that AI is more considerate, and thus, can enhance the intelligence of the smart life AI interaction system based on language recognition to improve the efficiency of elderly care.
[0115] In a specific implementation, when the first similarity is less than or equal to the second preset threshold, different physiological state parameters require different response roles. In a specific implementation, the mapping relationship between the preset physiological state parameters and the response roles can be pre-stored, and then the response role corresponding to the physiological state parameter can be determined based on the mapping relationship. The target language recognition system in the target area can use the response role to prompt the target object to re-enter the voice content.
[0116] Furthermore, the reply process corresponding to the reply role can also be recorded to generate a corresponding video file, and the video file can be sent to the reply role.
[0117] Optionally, the above step 204, adjusting the model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain target model parameters, may include the following steps:
[0118] Determining second initial model parameters corresponding to the target second language recognition model corresponding to the target identity information;
[0119] determining a first adjustment parameter corresponding to the target physiological state parameter;
[0120] The second initial model parameters are adjusted according to the first adjustment parameters to obtain the target second model parameters.
[0121] In the embodiment of the present application, different identity information requires different model parameters. Therefore, a mapping relationship between preset identity information and model parameters of a target second language recognition model can be pre-stored. Then, based on the mapping relationship, the second initial model parameters corresponding to the target second language recognition model corresponding to the target identity information can be determined.
[0122] Next, the mapping relationship between the preset physiological state parameters and the adjustment parameters can be pre-stored, and then, based on the mapping relationship, the first adjustment parameter corresponding to the target physiological state parameter can be determined, and then the second initial model parameter can be adjusted according to the first adjustment parameter to obtain the target second model parameter, that is, the target second model parameter = (1 + first adjustment parameter) * second initial model parameter. On the one hand, the corresponding model parameters can be configured based on the identity information of the target object, so that the model has an understanding ability that is adapted to the target object, avoiding understanding deviations during the AI interaction process. On the other hand, the target physiological state parameters reflect the actual needs and emotional changes of the target object to a certain extent, and the model parameters can also be dynamically adjusted based on the actual needs and emotional changes of the user. Therefore, it can not only avoid understanding deviations during the AI interaction process, but also ensure that the AI is more considerate, and then, it can improve the intelligence of the smart life AI interaction system based on language recognition to improve the efficiency of elderly care.
[0123] Optionally, the following steps may also be included:
[0124] Determining target sample data corresponding to the target area and the target identity information;
[0125] Using the target sample data to train the first language recognition model to obtain a trained first language recognition model;
[0126] The performing language recognition on the first speech signal according to the first language recognition model to obtain second text content includes:
[0127] determining first initial model parameters corresponding to the first speech recognition model corresponding to the target physiological state parameter;
[0128] determining a first signal-to-noise ratio relative to the first speech signal;
[0129] determining a second adjustment parameter corresponding to the first signal-to-noise ratio;
[0130] Adjusting the first initial model parameters according to the second adjustment parameters to obtain the target first model parameters;
[0131] Perform language recognition on the first speech signal according to the trained first language recognition model and the target first model parameters to obtain the second text content.
[0132] In a specific implementation, the target sample data corresponding to the target area and target identity information can be determined. The target sample data may include at least one of the following: the target object's self-voice data in the target area, the target object's conversational voice data in the target area, and the target object's AI interactive voice data in the target area. Usually, it can be understood as a period of historical time, which can be pre-set or system default.
[0133] In a specific implementation, the first language recognition model can be trained using target sample data related to the target region and the identity of the target object, and the trained first language recognition model is obtained, which is equivalent to upgrading the model and also has regional characteristics.
[0134] Next, a mapping relationship between preset physiological state parameters and model parameters corresponding to the first language recognition model may be pre-stored, and then, based on the mapping relationship, first initial model parameters corresponding to the first language recognition model corresponding to the target physiological state parameters may be determined.
[0135] Next, a first signal-to-noise ratio (SNR) of the first voice signal can be determined, and a mapping relationship between a preset SNR and an adjustment parameter can be pre-stored. Furthermore, a second adjustment parameter corresponding to the first SNR can be determined based on the mapping relationship. The first initial model parameter is then adjusted according to the second adjustment parameter to obtain a target first model parameter, where the target first model parameter = (1 + second adjustment parameter) * first initial model parameter. Language recognition is then performed on the first voice signal based on the trained first language recognition model and the target first model parameter to obtain the second text content. On the one hand, the first language recognition model can be trained using target sample data related to the target region and the identity of the target object to obtain a trained first language recognition model, which is equivalent to upgrading the model and also having regional characteristics. On the other hand, the target physiological state parameters reflect the actual needs and emotional changes of the target object to a certain extent. Not only can model parameters corresponding to the actual needs and emotional changes of the target object be obtained, but the model parameters can also be dynamically adjusted based on the SNR, making the model adaptive and anti-interference. This not only avoids misunderstandings during AI interaction, but also ensures that AI is more understanding. Furthermore, the intelligence of the smart life AI interaction system based on language recognition can be improved to improve the efficiency of elderly care.
[0136] Optionally, when the target area includes multiple microphones, the above step 202 of monitoring the first speech signal of the target object by the target language recognition system in the target area may include the following steps:
[0137] determining a sound source position by using the at least one camera of the target language recognition system in the target area, wherein the sound source position is a position of the mouth of the target object;
[0138] Collecting voice signals through the multiple microphones to obtain multiple voice signals;
[0139] Obtaining a position of each microphone in the plurality of microphones to obtain a plurality of positions;
[0140] determining a distance between each of the plurality of positions and the sound source position to obtain a plurality of distances;
[0141] performing noise reduction on corresponding voice signals among the plurality of voice signals according to the plurality of distances to obtain a plurality of noise-reduced voice signals;
[0142] The first speech signal is determined according to the multiple noise reduction speech signals.
[0143] In a specific implementation, the sound source position can be determined by at least one camera of the target language recognition system in the target area, where the sound source position is the position of the mouth of the target object. The voice signal can also be collected by multiple microphones to obtain multiple voice signals. The multiple microphones collect at the same time, or the collection time difference is less than the preset time length. The preset time length can be pre-set or the system defaults. In addition, the position of each microphone in the multiple microphones is obtained to obtain multiple positions. Since the position of each microphone is fixed, the distance between each position in the multiple positions and the sound source position can be determined to obtain multiple distances. The corresponding voice signals in the multiple voice signals are denoised according to the multiple distances to obtain multiple noise-reduced voice signals. Then, the first voice signal is determined based on the multiple noise-reduced voice signals. The noise can be dynamically reduced according to the actual distance of the sound transmission, which helps to ensure the noise reduction effect, and further, helps to improve the AI comprehension ability, and ensure the AI interaction ability and quality of smart life.
[0144] Optionally, the above step of performing noise reduction on corresponding voice signals among the multiple voice signals according to the multiple distances to obtain multiple noise-reduced voice signals may include the following steps:
[0145] Acquire first device information of the first microphone; the first microphone is any microphone among the multiple microphones;
[0146] determining a first noise reduction algorithm according to the first device information;
[0147] Acquire first algorithm control information of the first noise reduction algorithm;
[0148] determining first optimization information according to the distance corresponding to the first microphone;
[0149] determining target first algorithm control information according to the first optimization information and the first algorithm control information;
[0150] The speech signal corresponding to the first microphone is subjected to noise reduction using the target first algorithm control information and the first noise reduction algorithm to obtain a corresponding noise-reduced speech signal.
[0151] The first device information may include at least one of the following: the operating voltage of the microphone, the operating mode of the microphone, the model of the microphone, the operating current of the microphone, the operating power of the microphone, etc., which are not limited here.
[0152] In a specific implementation, taking the first microphone as an example, the first microphone is any one of multiple microphones, then the first device information of the first microphone can be obtained, and the mapping relationship between the preset device information and the noise reduction algorithm can be pre-stored, and then, the first noise reduction algorithm corresponding to the first device information can be determined based on the mapping relationship. In this way, a noise reduction algorithm that is deeply related to the performance of the device itself can be obtained.
[0153] The first algorithm control information may refer to the default algorithm control information of the first noise reduction algorithm, for example, it may control the algorithm speed, noise reduction capability, noise reduction frequency band, etc., which is not limited here. Next, the first algorithm control information of the first noise reduction algorithm may be obtained, and a mapping relationship between a preset distance and optimization information may be pre-stored. Then, the first optimization information corresponding to the distance corresponding to the first microphone may be determined based on the mapping relationship, and the target first algorithm control information may be determined based on the first optimization information and the first algorithm control information, where the target first algorithm control information = (1 + first optimization information) * first algorithm control information. Finally, the voice signal corresponding to the first microphone may be noise-reduced using the target first algorithm control information and the first noise reduction algorithm to obtain a corresponding noise-reduced voice signal. Furthermore, not only can the corresponding noise reduction algorithm be adaptively selected based on the actual performance (device parameters) of the microphone, but the algorithm control information of the noise reduction algorithm may also be dynamically optimized based on the distance, so that the depth of the noise reduction effect conforms to the actual spatial environment, thereby achieving deep noise reduction. In this way, dynamic noise reduction can be performed based on the actual distance of sound transmission, which helps to ensure the noise reduction effect, thereby helping to improve the AI comprehension ability, ensure the AI interaction capability of smart life, and ensure the quality of smart life.
[0154] Optionally, when the multiple microphones include m microphones, the above step of determining the first speech signal according to the multiple noise reduction speech signals may include the following steps:
[0155] Determining a signal-to-noise ratio of each of the m noise-reduced speech signals to obtain m signal-to-noise ratios;
[0156] Select a signal-to-noise ratio greater than the set signal-to-noise ratio from the m signal-to-noise ratios to obtain n target signal-to-noise ratios; n is a positive integer less than or equal to m;
[0157] Acquire speech signals corresponding to the n target signal-to-noise ratios to obtain n speech signals;
[0158] Determining a low-frequency component and a high-frequency component of each of the n speech signals;
[0159] Determining low-frequency energy proportions in the n speech signals to obtain n low-frequency energy proportions;
[0160] Determining high-frequency energy proportions in the n speech signals to obtain n high-frequency energy proportions;
[0161] Selecting low-frequency energy proportions within a preset range from the n low-frequency energy proportions to obtain a low-frequency energy proportions, where a is a positive integer less than or equal to n;
[0162] Determine the target low-frequency component according to the low-frequency components of the speech signal corresponding to the a low-frequency energy proportions;
[0163] Selecting the n high-frequency energy proportions whose high-frequency energy proportions are greater than the preset high-frequency energy proportion to obtain b high-frequency energy proportions, where b is a positive integer less than or equal to n;
[0164] Extracting features of the high-frequency components of the speech signal corresponding to the b high-frequency energy proportions to obtain b features;
[0165] Determine b weights according to the b features;
[0166] Determine a target high-frequency component according to the high-frequency components of the speech signal corresponding to the b high-frequency energy proportions and the b weights;
[0167] The first speech signal is determined according to the target low-frequency component and the target high-frequency component.
[0168] The signal-to-noise ratio may be preset or set by system default. The preset range may also be preset or set by system default. The preset high-frequency energy ratio may also be preset or set by system default.
[0169] In a specific implementation, the signal-to-noise ratio of each of the m noise-reduced speech signals can be determined to obtain m signal-to-noise ratios, and then a signal-to-noise ratio greater than a set signal-to-noise ratio is selected from the m signal-to-noise ratios to obtain n target signal-to-noise ratios, where n is a positive integer less than or equal to m. In this way, a high-quality speech signal can be selected for synthesizing the first speech signal.
[0170] The multi-scale decomposition algorithm may include at least one of the following: wavelet transform, pyramid transform, contourlet transform, etc., which are not limited here.
[0171] Next, speech signals corresponding to n target signal-to-noise ratios can be obtained to obtain n speech signals. Then, the low-frequency component and high-frequency component of each of the n speech signals are determined by a multi-scale decomposition algorithm. The low-frequency signal reflects the main part of the speech signal, and the high-frequency component reflects the detailed characteristics of the speech signal. Then, the low-frequency energy ratio in the n speech signals can be determined to obtain n low-frequency energy ratios, that is, low-frequency energy ratio = energy of low-frequency component / (energy of low-frequency component + energy of high-frequency component). Then, the high-frequency energy ratio in the n speech signals can be determined to obtain n high-frequency energy ratios, that is, high-frequency energy ratio = energy of high-frequency component / (energy of high-frequency component + energy of high-frequency component). Then, the low-frequency energy ratios in the n low-frequency energy ratios that are within a preset range can be selected from the n low-frequency energy ratios to obtain a low-frequency energy ratios, where a is a positive integer less than or equal to n. When the low-frequency energy ratio is greater than the upper limit of the preset range, it means that the main body masks the speech details. When the low-frequency energy ratio is less than the lower limit of the preset range, it means that the main body is poor, resulting in the details and the main body being mixed together. The low-frequency energy ratio within the preset range can ensure that the main body of the speech signal is full of quality.
[0172] Furthermore, the target low-frequency component can be determined based on the low-frequency components of the speech signal corresponding to the a low-frequency energy proportions. Specifically, the low-frequency components of the speech signal corresponding to the a low-frequency energy proportions can be averaged to obtain the target low-frequency component.
[0173] Furthermore, n high-frequency energy proportions greater than a preset high-frequency energy proportion can be selected to obtain b high-frequency energy proportions, where b is a positive integer less than or equal to n. In this way, a high-frequency part rich in details can be selected, and then the features of the high-frequency components of the speech signal corresponding to the b high-frequency energy proportions can be extracted to obtain b features. The features may include at least one of the following: feature points, feature densities, feature values, and feature vectors. Then, b weights can be determined based on the b features. For example, the greater the feature density, the greater the weight, and the greater the number of features, the greater the weight. In this way, the influence of the high-frequency components mixed with the selected details and the main body can be suppressed, weakening the influence of this part on the high frequency. , and then determine the target high-frequency component according to the high-frequency component of the speech signal corresponding to the b high-frequency energy proportions and b weights. Specifically, a weighted operation can be performed to obtain the high-frequency component. Finally, the first speech signal can be determined according to the target low-frequency component and the target high-frequency component. In this way, not only can a high-quality speech signal be selected to synthesize the first speech signal, but also a low-frequency component with full main body quality can be selected to synthesize the high-frequency component of the first speech signal. In addition, a high frequency with rich high-frequency details can be selected to synthesize the high-frequency component of the first speech signal while avoiding the influence of high-frequency components mixed with details and main body. In this way, a high-quality first speech signal can be obtained, which helps to ensure the comprehension ability of the model.
[0174] It can be seen that the language recognition-based smart life AI interaction method described in the embodiment of the present application is applied to a control platform device in a language recognition-based smart life AI interaction system, the control platform device includes a first language recognition model, the language recognition-based smart life AI interaction system also includes multiple language recognition systems, each language recognition system corresponds to an area and a second language recognition model, and the language recognition system of each area includes at least one positioning sensor, at least one camera, and at least one voice interaction device. The method includes: locating the target object through the at least one positioning sensor and the at least one camera in each area to determine the target area where the target object is located; monitoring the first voice signal of the target object through the target language recognition system in the target area; obtaining the target identity information and target physiological state parameters of the target object; adjusting the second model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain the target second model parameters; performing language recognition on the first voice signal through the target second language recognition model and the target second model parameters to obtain a first text content; performing language recognition on the first voice signal according to the first language recognition model to obtain a second text content; detecting a first similarity between the first text content and the second text content;When the first similarity is greater than a first preset threshold, a first response operation is performed according to the first text content. First, the target physiological state parameters reflect the actual needs and emotional changes of the target object to a certain extent. The second model parameters of the target second language recognition model of the target language recognition system can be adjusted according to the target identity information and the target physiological state parameters to obtain the target second model parameters. Obtaining model parameters that are deeply related to the actual situation of the target object helps to improve AI understanding capabilities and ensure the AI interaction capabilities and quality of smart life of smart life. Second, since the second language recognition model has regional characteristics, the recognition effect of the first text content also conforms to the regional characteristics. The first voice signal is language recognized by the target second language recognition model and the target second model parameters to obtain the first text content. The first text content has more regional characteristics. Third, since the first language recognition model can be based on the perspective of the target object's interactive object The training process also considers the language system (local characteristics) related to the target object. The first language recognition model is more locally specific and fully considers social (family and friends) understanding, not restricted by a single region. Fourthly, the first similarity between the first text content and the second text content is detected. To a certain extent, the similarity between the two includes the regional characteristics of the model, the local characteristics, and the full consideration of social (family and friends) understanding. When the first similarity is greater than a first preset threshold, it indicates that the model's understanding of the target object's language is deeply consistent with the regional characteristics and local characteristics, and fully considers social (family and friends) understanding. Furthermore, a first response operation can be executed based on the first text content. This not only avoids misunderstandings during AI interaction, but also ensures that AI is more understanding, thereby enhancing the intelligence of the smart life AI interaction system based on language recognition, thereby improving the efficiency of elderly care.
[0175] In accordance with the above embodiment, please refer to Figure 3 , Figure 3 This is a structural diagram of a control platform device provided in an embodiment of the present application. As shown in the figure, the control platform device includes a processor, a memory, a communication interface, and one or more programs. The one or more programs are stored in the memory and are configured to be executed by the processor. The control platform device is applied to a smart life AI interaction system based on language recognition. The control platform device includes a first language recognition model. The smart life AI interaction system based on language recognition also includes multiple language recognition systems, each language recognition system corresponds to a region and a second language recognition model. The language recognition system of each region includes at least one positioning sensor, at least one camera, and at least one voice interaction device. The program includes instructions for performing the following steps:
[0176] Positioning the target object by using the at least one positioning sensor and the at least one camera in each area to determine the target area where the target object is located;
[0177] monitoring a first speech signal of the target object by the target language recognition system in the target area;
[0178] Acquiring target identity information and target physiological state parameters of the target object;
[0179] adjusting second model parameters of a target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameter to obtain target second model parameters;
[0180] Performing language recognition on the first speech signal using the target second language recognition model and the target second model parameters to obtain first text content;
[0181] performing language recognition on the first speech signal according to the first language recognition model to obtain second text content;
[0182] detecting a first similarity between the first text content and the second text content;
[0183] When the first similarity is greater than a first preset threshold, a first response operation is performed according to the first text content.
[0184] Optionally, the program further includes instructions for executing the following steps:
[0185] When the first similarity is less than or equal to the first preset threshold and greater than a second preset threshold, acquiring, by the target language recognition system in the target area, video content of a preset time period corresponding to the first voice signal;
[0186] Performing action recognition on the video content to obtain third text content, wherein the third text content is used to represent the action intention or action state of the target object;
[0187] determining a second similarity between the first text content and the third text content;
[0188] determining a third similarity between the second text content and the third text content;
[0189] determining a larger value between the second similarity and the third similarity, and when the larger value is greater than a set threshold, performing a second response operation according to text content corresponding to the larger value;
[0190] When the larger value is less than or equal to the set threshold, the target language recognition system in the target area prompts the target subject to re-input voice content.
[0191] Optionally, the program further includes instructions for executing the following steps:
[0192] When the first similarity is less than or equal to the second preset threshold, the step of prompting the target object to re-input voice content through the target language recognition system in the target area is performed.
[0193] Optionally, in the aspect of adjusting the model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain target model parameters, the program includes instructions for executing the following steps:
[0194] Determining second initial model parameters corresponding to the target second language recognition model corresponding to the target identity information;
[0195] determining a first adjustment parameter corresponding to the target physiological state parameter;
[0196] The second initial model parameters are adjusted according to the first adjustment parameters to obtain the target second model parameters.
[0197] Optionally, the program further includes instructions for executing the following steps:
[0198] Determining target sample data corresponding to the target area and the target identity information;
[0199] Using the target sample data to train the first language recognition model to obtain a trained first language recognition model;
[0200] The performing language recognition on the first speech signal according to the first language recognition model to obtain second text content includes:
[0201] determining first initial model parameters corresponding to the first speech recognition model corresponding to the target physiological state parameter;
[0202] determining a first signal-to-noise ratio relative to the first speech signal;
[0203] determining a second adjustment parameter corresponding to the first signal-to-noise ratio;
[0204] Adjusting the first initial model parameters according to the second adjustment parameters to obtain the target first model parameters;
[0205] Perform language recognition on the first speech signal according to the trained first language recognition model and the target first model parameters to obtain the second text content.
[0206] It can be seen that the control platform device described in the embodiment of the present application is applied to a smart life AI interaction system based on language recognition. The control platform device includes a first language recognition model, and the smart life AI interaction system based on language recognition also includes multiple language recognition systems, each language recognition system corresponds to an area and a second language recognition model, and the language recognition system of each area includes at least one positioning sensor, at least one camera and at least one voice interaction device. The method includes: locating the target object through the at least one positioning sensor and the at least one camera in each area to determine the target area where the target object is located; monitoring the first voice signal of the target object through the target language recognition system in the target area; obtaining the target identity information and target physiological state parameters of the target object; adjusting the second model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain the target second model parameters; performing language recognition on the first voice signal through the target second language recognition model and the target second model parameters to obtain a first text content; performing language recognition on the first voice signal according to the first language recognition model to obtain a second text content; detecting a first similarity between the first text content and the second text content;When the first similarity is greater than a first preset threshold, a first response operation is performed according to the first text content. First, the target physiological state parameters reflect the actual needs and emotional changes of the target object to a certain extent. The second model parameters of the target second language recognition model of the target language recognition system can be adjusted according to the target identity information and the target physiological state parameters to obtain the target second model parameters. Obtaining model parameters that are deeply related to the actual situation of the target object helps to improve AI understanding capabilities and ensure the AI interaction capabilities and quality of smart life of smart life. Second, since the second language recognition model has regional characteristics, the recognition effect of the first text content also conforms to the regional characteristics. The first voice signal is language recognized by the target second language recognition model and the target second model parameters to obtain the first text content. The first text content has more regional characteristics. Third, since the first language recognition model can be based on the perspective of the target object's interactive object The training process also considers the language system (local characteristics) related to the target object. The first language recognition model is more locally specific and fully considers social (family and friends) understanding, not restricted by a single region. Fourthly, the first similarity between the first text content and the second text content is detected. To a certain extent, the similarity between the two includes the regional characteristics of the model, the local characteristics, and the full consideration of social (family and friends) understanding. When the first similarity is greater than a first preset threshold, it indicates that the model's understanding of the target object's language is deeply consistent with the regional characteristics and local characteristics, and fully considers social (family and friends) understanding. Furthermore, a first response operation can be executed based on the first text content. This not only avoids misunderstandings during AI interaction, but also ensures that AI is more understanding, thereby enhancing the intelligence of the smart life AI interaction system based on language recognition, thereby improving the efficiency of elderly care.
[0207] Figure 4 This is a functional unit block diagram of a language recognition-based smart life AI interaction system 400 involved in an embodiment of the present application. The language recognition-based smart life AI interaction system 400 is applied to a control platform device in the language recognition-based smart life AI interaction system. The control platform device includes a first language recognition model. The language recognition-based smart life AI interaction system also includes multiple language recognition systems, each language recognition system corresponding to a region and a second language recognition model. The language recognition system in each region includes at least one positioning sensor, at least one camera, and at least one voice interaction device. The language recognition-based smart life AI interaction system 400 includes:
[0208] A positioning unit 401 is configured to locate a target object by using the at least one positioning sensor and the at least one camera in each area to determine a target area where the target object is located;
[0209] A monitoring unit 402 is configured to monitor a first speech signal of the target object through the target language recognition system in the target area;
[0210] An acquisition unit 403 is configured to acquire target identity information and target physiological state parameters of the target object;
[0211] An adjusting unit 404 is configured to adjust a second model parameter of a target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameter to obtain a target second model parameter;
[0212] The recognition unit 405 is configured to perform language recognition on the first speech signal using the target second language recognition model and the target second model parameters to obtain first text content; and perform language recognition on the first speech signal according to the first language recognition model to obtain second text content;
[0213] A detection unit 406, configured to detect a first similarity between the first text content and the second text content;
[0214] The execution unit 407 is configured to execute a first response operation according to the first text content when the first similarity is greater than a first preset threshold.
[0215] Optionally, the language recognition-based smart life AI interaction system 400 is specifically used to:
[0216] When the first similarity is less than or equal to the first preset threshold and greater than a second preset threshold, acquiring, by the target language recognition system in the target area, video content of a preset time period corresponding to the first voice signal;
[0217] Performing action recognition on the video content to obtain third text content, wherein the third text content is used to represent the action intention or action state of the target object;
[0218] determining a second similarity between the first text content and the third text content;
[0219] determining a third similarity between the second text content and the third text content;
[0220] determining a larger value between the second similarity and the third similarity, and when the larger value is greater than a set threshold, performing a second response operation according to text content corresponding to the larger value;
[0221] When the larger value is less than or equal to the set threshold, the target language recognition system in the target area prompts the target subject to re-input voice content.
[0222] Optionally, the language recognition-based smart life AI interaction system 400 is specifically used to:
[0223] When the first similarity is less than or equal to the second preset threshold, the step of prompting the target object to re-input voice content through the target language recognition system in the target area is performed.
[0224] Optionally, in the aspect of adjusting the model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain target model parameters, the adjusting unit 404 is specifically configured to:
[0225] Determining second initial model parameters corresponding to the target second language recognition model corresponding to the target identity information;
[0226] determining a first adjustment parameter corresponding to the target physiological state parameter;
[0227] The second initial model parameters are adjusted according to the first adjustment parameters to obtain the target second model parameters.
[0228] Optionally, the language recognition-based smart life AI interaction system 400 is specifically used to:
[0229] Determining target sample data corresponding to the target area and the target identity information;
[0230] Using the target sample data to train the first language recognition model to obtain a trained first language recognition model;
[0231] The performing language recognition on the first speech signal according to the first language recognition model to obtain second text content includes:
[0232] determining first initial model parameters corresponding to the first speech recognition model corresponding to the target physiological state parameter;
[0233] determining a first signal-to-noise ratio relative to the first speech signal;
[0234] determining a second adjustment parameter corresponding to the first signal-to-noise ratio;
[0235] Adjusting the first initial model parameters according to the second adjustment parameters to obtain the target first model parameters;
[0236] Perform language recognition on the first speech signal according to the trained first language recognition model and the target first model parameters to obtain the second text content.
[0237] It can be seen that the language recognition-based smart life AI interaction system described in the embodiment of the present application is applied to a control platform device in the language recognition-based smart life AI interaction system, the control platform device includes a first language recognition model, the language recognition-based smart life AI interaction system also includes multiple language recognition systems, each language recognition system corresponds to an area and a second language recognition model, and the language recognition system of each area includes at least one positioning sensor, at least one camera and at least one voice interaction device. The method includes: locating the target object by using the at least one positioning sensor and the at least one camera in each area to determine the target area where the target object is located; monitoring the first voice signal of the target object by using the target language recognition system in the target area; obtaining the target identity information and target physiological state parameters of the target object; adjusting the second model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain the target second model parameters; performing language recognition on the first voice signal by using the target second language recognition model and the target second model parameters to obtain a first text content; performing language recognition on the first voice signal according to the first language recognition model to obtain a second text content; detecting a first similarity between the first text content and the second text content;When the first similarity is greater than a first preset threshold, a first response operation is performed according to the first text content. First, the target physiological state parameters reflect the actual needs and emotional changes of the target object to a certain extent. The second model parameters of the target second language recognition model of the target language recognition system can be adjusted according to the target identity information and the target physiological state parameters to obtain the target second model parameters. Obtaining model parameters that are deeply related to the actual situation of the target object helps to improve AI understanding capabilities and ensure the AI interaction capabilities and quality of smart life of smart life. Second, since the second language recognition model has regional characteristics, the recognition effect of the first text content also conforms to the regional characteristics. The first voice signal is language recognized by the target second language recognition model and the target second model parameters to obtain the first text content. The first text content has more regional characteristics. Third, since the first language recognition model can be based on the perspective of the target object's interactive object The training process also considers the language system (local characteristics) related to the target object. The first language recognition model is more locally specific and fully considers social (family and friends) understanding, not restricted by a single region. Fourthly, the first similarity between the first text content and the second text content is detected. To a certain extent, the similarity between the two includes the regional characteristics of the model, the local characteristics, and the full consideration of social (family and friends) understanding. When the first similarity is greater than a first preset threshold, it indicates that the model's understanding of the target object's language is deeply consistent with the regional characteristics and local characteristics, and fully considers social (family and friends) understanding. Furthermore, a first response operation can be executed based on the first text content. This not only avoids misunderstandings during AI interaction, but also ensures that AI is more understanding, thereby enhancing the intelligence of the smart life AI interaction system based on language recognition, thereby improving the efficiency of elderly care.
[0238] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute part or all of the steps of any method described in the above method embodiments.
[0239] The present application also provides a computer program product comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package.
[0240] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0241] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0242] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0243] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0244] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0245] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a memory, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the above-mentioned methods of each embodiment of the present application. The aforementioned memory includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0246] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable memory, and the memory can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0247] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A smart life AI interaction method based on language recognition, characterized in that: A control platform device applied to a language recognition-based smart life AI interaction system, the control platform device including a server; the control platform device including a first language recognition model, the first language recognition model including a large language model; the language recognition-based smart life AI interaction system also including multiple language recognition systems, each language recognition system corresponding to a region and a second language recognition model, the language recognition system in each region including at least one positioning sensor, at least one camera, and at least one voice interaction device, the method comprising: Positioning a target object by using the at least one positioning sensor and the at least one camera in each area to determine a target area where the target object is located; the target object includes an elderly person; monitoring a first speech signal of the target object by a target language recognition system in the target area; Acquiring target identity information and target physiological state parameters of the target object; the target physiological state parameters are used to reflect the actual needs and emotional changes of the target object; adjusting a second model parameter of a target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameter to obtain a target second model parameter; wherein the target second language recognition model has regional characteristics; Performing language recognition on the first speech signal using the target second language recognition model and the target second model parameters to obtain first text content; performing language recognition on the first speech signal according to the first language recognition model to obtain second text content; the first language recognition model is trained based on the perspective of the target object's interactive object and takes into account the language system related to the target object; detecting a first similarity between the first text content and the second text content; When the first similarity is greater than a first preset threshold, a first response operation is performed according to the first text content.
2. The method according to claim 1, characterized in that The method further comprises: When the first similarity is less than or equal to the first preset threshold and greater than a second preset threshold, acquiring, by the target language recognition system in the target area, video content of a preset time period corresponding to the first voice signal; Performing action recognition on the video content to obtain third text content, wherein the third text content is used to represent the action intention or action state of the target object; determining a second similarity between the first text content and the third text content; determining a third similarity between the second text content and the third text content; determining a larger value between the second similarity and the third similarity, and when the larger value is greater than a set threshold, performing a second response operation according to text content corresponding to the larger value; When the larger value is less than or equal to the set threshold, the target language recognition system in the target area prompts the target subject to re-input voice content.
3. The method according to claim 2, characterized in that The method further comprises: When the first similarity is less than or equal to the second preset threshold, the step of prompting the target object to re-input voice content through the target language recognition system in the target area is performed.
4. The method according to any one of claims 1 to 3, characterized in that The step of adjusting the model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain target model parameters includes: Determining second initial model parameters corresponding to the target second language recognition model corresponding to the target identity information; determining a first adjustment parameter corresponding to the target physiological state parameter; The second initial model parameters are adjusted according to the first adjustment parameters to obtain the target second model parameters.
5. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Determining target sample data corresponding to the target area and the target identity information; Using the target sample data to train the first language recognition model to obtain a trained first language recognition model; The performing language recognition on the first speech signal according to the first language recognition model to obtain second text content includes: determining first initial model parameters corresponding to the first speech recognition model corresponding to the target physiological state parameter; determining a first signal-to-noise ratio relative to the first speech signal; determining a second adjustment parameter corresponding to the first signal-to-noise ratio; Adjusting the first initial model parameters according to the second adjustment parameters to obtain target first model parameters; Perform language recognition on the first speech signal according to the trained first language recognition model and the target first model parameters to obtain the second text content.
6. A smart life AI interaction system based on language recognition, characterized by: A control platform device used in the language recognition-based smart life AI interaction system, the control platform device including a server; the control platform device including a first language recognition model, the first language recognition model including a large language model; the language recognition-based smart life AI interaction system also including multiple language recognition systems, each language recognition system corresponding to a region and a second language recognition model, the language recognition system in each region including at least one positioning sensor, at least one camera, and at least one voice interaction device, the language recognition-based smart life AI interaction system including: a positioning unit, configured to locate a target object by using the at least one positioning sensor and the at least one camera in each area to determine a target area where the target object is located; the target object includes an elderly person; A monitoring unit, configured to monitor a first speech signal of the target object through a target language recognition system in the target area; an acquisition unit, configured to acquire target identity information and target physiological state parameters of the target object; the target physiological state parameters are used to reflect the actual needs and emotional changes of the target object; an adjusting unit, configured to adjust a second model parameter of a target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameter to obtain a target second model parameter; the target second language recognition model having regional characteristics; a recognition unit configured to perform language recognition on the first speech signal using the target second language recognition model and the target second model parameters to obtain first text content; and perform language recognition on the first speech signal based on the first language recognition model to obtain second text content; the first language recognition model is trained based on the perspective of an interactive object of the target object and takes into account a language system related to the target object; a detection unit, configured to detect a first similarity between the first text content and the second text content; An execution unit is configured to execute a first response operation according to the first text content when the first similarity is greater than a first preset threshold.
7. The intelligent life AI interactive system based on language recognition according to claim 6, characterized in that: The language recognition-based smart life AI interaction system is specifically used for: When the first similarity is less than or equal to the first preset threshold and greater than a second preset threshold, acquiring, by the target language recognition system in the target area, video content of a preset time period corresponding to the first voice signal; Performing action recognition on the video content to obtain third text content, wherein the third text content is used to represent the action intention or action state of the target object; determining a second similarity between the first text content and the third text content; determining a third similarity between the second text content and the third text content; determining a larger value between the second similarity and the third similarity, and when the larger value is greater than a set threshold, performing a second response operation according to text content corresponding to the larger value; When the larger value is less than or equal to the set threshold, the target language recognition system in the target area prompts the target subject to re-input voice content.
8. The intelligent life AI interactive system based on language recognition according to claim 7 is characterized in that: The language recognition-based smart life AI interaction system is also specifically used for: When the first similarity is less than or equal to the second preset threshold, the step of prompting the target object to re-input voice content through the target language recognition system in the target area is performed.
9. The intelligent life AI interactive system based on language recognition according to any one of claims 6 to 8, characterized in that: In the aspect of adjusting the model parameters of the target second language recognition model of the target language recognition system according to the target identity information and the target physiological state parameters to obtain the target model parameters, the adjusting unit is specifically configured to: Determining second initial model parameters corresponding to the target second language recognition model corresponding to the target identity information; determining a first adjustment parameter corresponding to the target physiological state parameter; The second initial model parameters are adjusted according to the first adjustment parameters to obtain the target second model parameters.
10. The intelligent life AI interactive system based on language recognition according to any one of claims 6 to 8, characterized in that: The language recognition-based smart life AI interaction system is also specifically used for: Determining target sample data corresponding to the target area and the target identity information; Using the target sample data to train the first language recognition model to obtain a trained first language recognition model; The performing language recognition on the first speech signal according to the first language recognition model to obtain second text content includes: determining first initial model parameters corresponding to the first speech recognition model corresponding to the target physiological state parameter; determining a first signal-to-noise ratio relative to the first speech signal; determining a second adjustment parameter corresponding to the first signal-to-noise ratio; Adjusting the first initial model parameters according to the second adjustment parameters to obtain target first model parameters; Perform language recognition on the first speech signal according to the trained first language recognition model and the target first model parameters to obtain the second text content.
Citation Information
Patent Citations
Speech recognition models based on location indicia
CN104509079A
Intelligent voice recognition method based on three-level feature collection
CN111986674A