Intelligent voice interaction method for old people based on large model

By collecting voice data from elderly users to generate user profiles and utilizing large models and dialect databases, the problem of mismatch between smart devices and the elderly has been solved, enabling personalized interactive data push and improving the ease of use for the elderly and the applicability of the devices.

CN121053979APending Publication Date: 2025-12-02TIANHE COLLEGE GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510973984.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing smart devices have problems when interacting with the elderly, such as difficulty in matching dialects, complex interaction processes, and inconvenient operation, which fail to meet the needs of the elderly and limit the scope of application of the devices.

Method used

By collecting voice data from elderly users, determining user text data and generating user profiles, and using large models and dialect databases to generate personalized interactive push data and target voice data, the matching of voice output with the user's language environment is improved.

Benefits of technology

It enables personalized interactive data push, improves the ease of use for elderly users, and expands the applicability of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053979A_ABST
    Figure CN121053979A_ABST
Patent Text Reader

Abstract

The invention discloses an old people intelligent voice interaction method based on a large model, and belongs to the technical field of intelligent equipment. The method comprises the following steps: acquiring user voice data of an elderly user, and determining user text data according to the user voice data; determining a user portrait of the elderly user according to the user text data and a preset large model; determining interactive push data according to the user portrait data; and generating target voice data according to the interactive push data and the dialect database. The beneficial effect of improving the application range of the device is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart devices, and in particular to a smart voice interaction method for the elderly based on a large model. Background Technology

[0002] With the widespread use of smart devices in homes, healthcare, transportation, and other fields, the elderly are also gradually using them. Currently, smart devices are often designed for the needs of the general population, and their adaptability to the elderly is insufficient. Common issues include: the elderly having difficulty interacting with smart devices due to their dialects, complex interaction processes, lengthy instructions that do not match the cognitive characteristics and operating habits of the elderly, and a lack of interactive content tailored to the needs of the elderly. As a result, smart devices cannot effectively serve the elderly, thus reducing the applicability of the devices.

[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this invention is to provide a method, device, smart equipment, and storage medium for intelligent voice interaction for the elderly based on a large model, aiming to improve the applicability of the device.

[0005] To achieve the above objectives, the present invention provides a large-scale model-based intelligent voice interaction method for the elderly, which includes the following steps:

[0006] Collect user voice data from elderly users and determine user text data based on the user voice data;

[0007] The user profile of the elderly user is determined based on the user text data and the preset large model;

[0008] The interactive push data is determined based on the user profile data;

[0009] Target speech data is generated based on the interactive push data and the dialect database.

[0010] Optionally, the number of user text data is multiple, and the step of determining the user profile of the elderly user based on the user text data and a preset large model includes:

[0011] Based on the text time data and the first preset profile, multiple user text data are combined to obtain target text data, where the text time data is the time data corresponding to the user text data.

[0012] First input data is generated based on the preset first portrait extraction instruction and the target text data;

[0013] The first input data is input into the preset large model to obtain the model output, and the user profile is determined based on the model output.

[0014] Optionally, the text time data includes: text start time, text end time, and text duration. The step of combining multiple user text data according to the text time data and the first preset profile to obtain target text data includes:

[0015] The text duration difference is obtained by comparing the time difference between two adjacent user text data and comparing the text duration of the two adjacent user text data.

[0016] Based on the time difference and the text duration difference, determine the text time characteristics of two time-adjacent user text data;

[0017] Based on the text time characteristics and the first preset profile, the multiple user text data are combined to obtain target text data;

[0018] The time difference is the difference between the start time of the later user text data and the end time of the earlier user text data.

[0019] Optionally, the step of combining the plurality of user text data according to the text time features and the first preset profile to obtain target text data includes:

[0020] When the text time feature is a first type of time feature, and one of the two user text data that are adjacent in time matches the portrait data of the first preset portrait, the user text data that does not match the portrait data of the first preset portrait in the two user text data that are adjacent in time is deleted.

[0021] When the text time feature is a feature other than the first type of time feature, or when two user text data that are adjacent in time both match the portrait data of the first preset portrait, the two user text data that are adjacent in time are combined, and the combined text data is used as the updated user text data.

[0022] When the number of user text data is 1, the user text data is used as the target text data;

[0023] The first type of time feature is that the time difference is less than or equal to a preset time difference, and the text duration difference is less than or equal to a preset text duration difference.

[0024] Optionally, after the step of inputting the first input data into the preset large model to obtain the model output, and determining the user profile based on the model output, the method further includes:

[0025] Calculate the cosine similarity between the user profile and the first preset profile;

[0026] When the cosine similarity is less than or equal to the preset similarity, the first preset profile is updated according to the user profile;

[0027] When the cosine similarity is greater than the preset similarity, the first preset portrait is updated according to the portrait database.

[0028] Optionally, the step of determining the interactive push data based on the user profile data includes:

[0029] The first demand information is determined based on the user profile data, and the second demand information is determined based on the user text data. The first demand information is emotional demand information, and the second demand information is the actual demand information at the current moment.

[0030] Interactive push data is generated based on the current scene data, the primary demand information, and the secondary demand information.

[0031] Optionally, the step of generating target speech data based on the interactive push data and the dialect database includes:

[0032] The interactive push data is standardized to obtain standard speech conversion data;

[0033] Prosody prediction is performed based on the standard speech conversion data to obtain speech feature data of the speech data. The speech feature data includes: the pause position, stress, intonation, and sentence duration of each sentence in the standard speech conversion data.

[0034] The target speech data is generated based on the speech feature data, the pronunciation dictionary, and the standard speech conversion data.

[0035] Furthermore, to achieve the above objectives, the present invention also provides a large-scale intelligent voice interaction device for the elderly, the large-scale intelligent voice interaction device for the elderly comprising:

[0036] The acquisition module is used to collect user voice data of elderly users and determine user text data based on the user voice data;

[0037] The identification module is used to determine the user profile of the elderly user based on the user text data and a preset large model;

[0038] The push module is used to determine interactive push data based on the user profile data;

[0039] The voice module is used to generate target voice data based on the interactive push data and the dialect database.

[0040] Furthermore, to achieve the above objectives, the present invention also provides an intelligent device, the intelligent device comprising: a memory, a processor, and a large-model-based intelligent voice interaction program for the elderly stored in the memory and executable on the processor, the large-model-based intelligent voice interaction program for the elderly being configured to implement the steps of the large-model-based intelligent voice interaction method for the elderly as described above.

[0041] In addition, to achieve the above objectives, the present invention also provides a storage medium storing a large-model-based intelligent voice interaction program for the elderly, wherein when the large-model-based intelligent voice interaction program for the elderly is executed by a processor, it implements the steps of the large-model-based intelligent voice interaction method for the elderly described above.

[0042] This invention proposes a large-scale model-based intelligent voice interaction method for the elderly. It collects voice data from elderly users, determines text data based on this voice data, and then uses this text data and a pre-set large-scale model to create a user profile for the elderly user. Compared to current single-method information push, this method determines interactive push data based on the user profile data, enabling accurate and personalized interactive data push based on the identified elderly user. Furthermore, it generates target voice data based on the interactive push data and a dialect database, improving the matching of the output voice with the user's actual language environment, thereby enhancing the ease of use for elderly users and expanding the applicability of the device. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of the structure of an intelligent device in the hardware operating environment involved in the embodiments of the present invention;

[0044] Figure 2 This is a flowchart illustrating the first embodiment of the intelligent voice interaction method for the elderly based on a large model according to the present invention.

[0045] Figure 3 This is a flowchart illustrating the second embodiment of the intelligent voice interaction method for the elderly based on a large model according to the present invention.

[0046] Figure 4 This is a flowchart illustrating the third embodiment of the intelligent voice interaction method for the elderly based on a large model according to the present invention.

[0047] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0048] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0049] Reference Figure 1 , Figure 1 This is a schematic diagram of the intelligent device structure of the hardware operating environment involved in the embodiments of the present invention.

[0050] like Figure 1 As shown, the smart device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, an interaction device 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The interaction device 1003 may include a display screen and an input unit such as a keyboard. Optionally, the interaction device 1003 may also be connected to the communication bus via a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0051] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on smart devices and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0052] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and a smart voice interaction program for the elderly based on a large model.

[0053] exist Figure 1 In the smart device shown, the network interface 1004 is mainly used for data communication with other devices; the interaction device 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the smart device of the present invention can be set in the smart device. The smart device calls the elderly intelligent voice interaction program based on a large model stored in the memory 1005 through the processor 1001 and executes the elderly intelligent voice interaction method based on a large model provided in the embodiment of the present invention.

[0054] This invention provides a method for intelligent voice interaction for the elderly based on a large model, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of a method for intelligent voice interaction for the elderly based on a large model according to the present invention.

[0055] In this embodiment, the intelligent voice interaction method for the elderly based on a large model includes:

[0056] Step S1: Collect user voice data of elderly users and determine user text data based on the user voice data;

[0057] In this embodiment, a microphone is provided, preferably a high-sensitivity microphone, to ensure that even if the elderly user's speech is unclear, it can still be collected. Specifically, after collecting the speech data, the user's speech data is identified using speech recognition technology and a dialect database, and then converted into user text data.

[0058] Step S2: Determine the user profile of the elderly user based on the user text data and the preset large model;

[0059] In this embodiment, a corresponding user profile extraction instruction is set. Specifically, since user profiles are generally stored in a structured manner, i.e., multiple different tags are set, the vector data composed of these tags constitutes the user profile. Each element of the vector stores a type of tag, and these tags can differ between different vectors, corresponding to different user profiles. Specifically, the user profile extraction instruction includes a feature table for the user profile, which corresponds to the vector data composed of the tags. The user profile extraction instruction is input simultaneously with the user text data, enabling the pre-defined large model to extract user profile-related data according to the format requirements of the user profile extraction instruction and output the results according to the requirements of the user profile feature table.

[0060] Step S3: Determine interactive push data based on the user profile data;

[0061] In this embodiment, targeted data can be pushed based on the current user profile. Optionally, one type of user profile data can correspond to one type of interactive push data. Optionally, one type of user profile data can correspond to more than one type of interactive push data, wherein different types of interactive push data have different push probabilities. Optionally, based on the user profile and historical behavior data accumulated before the current moment, the system analyzes the user's personalized needs and determines the content required by the user. Commonly, interactive push data can be: for users with health management needs, the system pushes medication reminders regularly based on their entered health information and medical orders, and links to a professional health consultation platform to provide authoritative health knowledge answers.

[0062] Step S4: Generate target speech data based on the interactive push data and the dialect database.

[0063] Optionally, the interactive push data can be processed using speech synthesis technology and a dialect database to generate corresponding target speech data. Alternatively, pre-recorded dialect speech segments can be used for output. While this method reduces technical complexity, specifically, the interactive push data is converted into dialect text data with the same meaning or type. Specifically, semantic matching can be used to determine dialect text data with similar meanings to the interactive push data, and the target speech data can be generated by synthesizing the dialect text data and at least one pre-recorded dialect speech segment.

[0064] In this embodiment, by collecting user voice data from elderly users and determining user text data based on the voice data, and then determining the user profile of the elderly user based on the user text data and a preset large model, compared with the current single information push method, the interactive push data determined based on the user profile data can accurately achieve personalized interactive data push based on the identification of the elderly. Furthermore, by generating target voice data based on the interactive push data and a dialect database, the output voice can be better matched with the user's actual language environment, thereby improving the convenience of use for elderly users and increasing the applicability of the device.

[0065] Furthermore, based on the first embodiment, a second embodiment of the present invention's intelligent voice interaction method for the elderly based on a large model is proposed. In this embodiment, reference is made to... Figure 3 The number of user text data is multiple, and the step of determining the user profile of the elderly user based on the user text data and the preset large model includes:

[0066] Step S21: Combine multiple user text data according to the text time data and the first preset profile to obtain target text data, wherein the text time data is the time data corresponding to the user text data;

[0067] It should be noted that elderly users interact with the device according to their own usage habits during actual communication, such as pausing for a period of time before interacting with the device. Therefore, multiple sets of user text data may exist. When using smart devices, elderly people often repeat input due to inaccurate pronunciation; however, only one of these repeated inputs may be correctly recognized and meet the needs of the elderly user. Therefore, in combining multiple sets of user text data, it is necessary to combine text time data and a first preset profile to determine which are repetitive.

[0068] Step S22: Generate first input data according to the preset first portrait extraction instruction and the target text data;

[0069] After obtaining the target text data, the preset first portrait extraction instruction and the target text data are fused together to obtain the first input data. This is to enable the complete data instruction and data to be input into the preset large model in one go.

[0070] Step S23: Input the first input data into the preset large model to obtain the model output, and determine the user profile based on the model output.

[0071] The first input data is input into the preset large model to obtain the model output of the preset large model based on the first input data. The user profile is then determined based on the model output. Here, the preset large model can be a general-purpose large language model, such as DeepSeek. Preferably, a usable database for the large model can be set simultaneously with the preset large model. The large model can complete the user profile recognition based on the information in the usable database. The usable database can be a database containing user profile recognition terms.

[0072] In this embodiment, multiple user text data are combined using text time data and a first preset profile to obtain target text data. The text time data is the time data corresponding to the user text data. First input data is generated according to the preset first profile extraction instruction and the target text data. The first input data is input into the preset large model to obtain the model output. The user profile is determined according to the model output, thereby enabling the identification of user profiles with complete user information and improving the accuracy of identification.

[0073] Furthermore, based on the first or second embodiment, a third embodiment of the present invention for the intelligent voice interaction method for the elderly based on a large model is proposed. In this embodiment, reference is made to... Figure 4The text time data includes: text start time, text end time, and text duration. The step of combining multiple user text data according to the text time data and the first preset profile to obtain target text data includes:

[0074] Step S231: Based on the time difference between two adjacent user text data, and by comparing the text duration of the two adjacent user text data, obtain the text duration difference;

[0075] Currently, because elderly users find it difficult to manually delete incorrectly identified text data using complex methods, multiple erroneous text data entries may exist among the obtained user text data. Furthermore, relying solely on semantic analysis for identification often fails to recognize these erroneous text entries. Therefore, in this embodiment, the time difference and duration difference between two adjacent user text data entries are extracted to determine whether they exhibit temporal repetition of the same utterance. In this embodiment, temporal adjacency means that within the two time ranges corresponding to the two collected user text data entries, there is no other user text data. Therefore, this embodiment does not limit the range of the time difference.

[0076] Step S232: Determine the text time characteristics of two adjacent user text data based on the time difference and the text duration difference;

[0077] Optionally, the time difference and the text duration difference are concatenated to generate the text time vector as a text time feature.

[0078] Step S233: Combine the multiple user text data according to the text time features and the first preset profile to obtain target text data;

[0079] The first preset profile here is actually created by deleting data that does not match the profile during the process of combining multiple user text data based on the user's historical data.

[0080] The time difference is the difference between the start time of the later user text data and the end time of the earlier user text data.

[0081] In this embodiment, the text duration difference is obtained by comparing the time difference between two adjacent user text data and comparing the text duration of the two adjacent user text data. The text time characteristics of the two adjacent user text data are determined based on the time difference and the text duration difference. In fact, the time characteristics can be used to preferentially filter user text data with a high probability of being repeated input, thereby avoiding the analysis of all user text data through the first preset profile, improving the efficiency of identifying abnormal data. The multiple user text data are combined according to the text time characteristics and the first preset profile to obtain target text data, thereby improving the accuracy of the combined data while combining the multiple user text data.

[0082] Furthermore, the step of combining the multiple user text data according to the text time features and the first preset profile to obtain the target text data includes:

[0083] When the text time feature is a first type of time feature, and one of the two user text data that are adjacent in time matches the portrait data of the first preset portrait, the user text data that does not match the portrait data of the first preset portrait in the two user text data that are adjacent in time is deleted.

[0084] When the text time feature is a feature other than the first type of time feature, or when two user text data that are adjacent in time both match the portrait data of the first preset portrait, the two user text data that are adjacent in time are combined, and the combined text data is used as the updated user text data.

[0085] When the number of user text data is 1, the user text data is used as the target text data;

[0086] The first type of time feature is that the time difference is less than or equal to a preset time difference, and the text duration difference is less than or equal to a preset text duration difference.

[0087] Optionally, both the preset time difference and the preset text duration difference can be 2 seconds, which can be set by the manufacturer. Furthermore, when the first type of time feature is present, potentially abnormal data can be identified. For further identification, if one of two adjacent user text data sets matches the profile data of the first preset profile, it is determined that there is duplicate abnormal user text data, and the user text data that does not match the profile data of the first preset profile among the two adjacent user text data sets is deleted.

[0088] Furthermore, it should be noted that when the text time feature is a first type of time feature, and no user text data in two adjacent user text data sets matches the profile data of the first preset profile, one of the user text data sets is deleted. Specifically, the specific user text data to be deleted is determined through semantic analysis. Thus, in this embodiment, even if an elderly user repeatedly inputs the same meaning more than once, the duplicate user text data with recognition anomalies can be deleted. In other embodiments, when the number of user text data sets is a preset number of user text data sets, the user text data sets are used as the target text data sets. Commonly, associated vocabulary data is determined based on the profile data of the first preset profile, semantic similarity is calculated between the associated vocabulary data sets and the user text data sets, and a match is determined based on the semantic similarity. The associated vocabulary data sets include multiple words, short phrases, etc.

[0089] In this embodiment, by combining the text time features corresponding to the user text data, the abnormal text time features can be effectively limited, and different text merging methods can be performed according to the matching results of the user text data and the portrait data of the first preset portrait, thereby improving the accuracy of the target text data obtained after merging.

[0090] Furthermore, based on any of the above embodiments, a fourth embodiment of the present invention for the intelligent voice interaction method for the elderly based on a large model is proposed. In this embodiment, after the steps of inputting the first input data into the preset large model to obtain the model output and determining the user profile based on the model output, the method further includes:

[0091] Calculate the cosine similarity between the user profile and the first preset profile;

[0092] When the cosine similarity is less than or equal to the preset similarity, the first preset profile is updated according to the user profile;

[0093] When the cosine similarity is greater than the preset similarity, the first preset portrait is updated according to the portrait database.

[0094] Specifically, by calculating the cosine similarity between the user profile and the first preset profile, it is determined whether the user profile of elderly users needs to be updated in real time. Even among elderly users, there can be different elderly users; for example, the preset profiles for male elderly users and female elderly users can be different. Optionally, the corresponding first preset profile is determined based on voiceprint information extracted from the user's voice data. For example, the first preset profile of a user already recorded can be derived from voiceprint information. When the cosine similarity is greater than the preset similarity, the first preset profile is updated according to the profile database, indicating an anomaly in the setting or selection of the first preset profile; therefore, the first preset profile needs to be updated according to the profile database.

[0095] In this embodiment, by calculating the cosine similarity between the user profile and the first preset profile, when the cosine similarity is less than or equal to the preset similarity, the first preset profile is updated according to the user profile; when the cosine similarity is greater than the preset similarity, the first preset profile is updated according to the profile database. This improves the accuracy of the preset profile and enables real-time updates of the user profile.

[0096] Furthermore, based on any of the above embodiments, a fifth embodiment of the present invention for the intelligent voice interaction method for the elderly based on a large model is proposed. In this embodiment, the step of determining the interaction push data based on the user profile data includes:

[0097] The first demand information is determined based on the user profile data, and the second demand information is determined based on the user text data. The first demand information is emotional demand information, and the second demand information is the actual demand information at the current moment.

[0098] Interactive push data is generated based on the current scene data, the primary demand information, and the secondary demand information.

[0099] Optionally, the interaction database is filtered based on the current scene data, the first demand information, and the second demand information. This database can be updated in real time, for example, with news and community service information, to obtain information such as news and community services needed by elderly users as the number of interaction pushes. In other embodiments, sentiment analysis can be implemented using rule-based sentiment analysis algorithms. Alternatively, the first demand information can be determined using a third-party sentiment analysis API, such as the Microsoft Azure Emotion API, which can effectively reduce costs. Furthermore, in other embodiments, services can also be provided based on user-defined service preferences.

[0100] In this embodiment, first demand information is determined based on the user profile data, and second demand information is determined based on the user text data. The first demand information is emotional demand information, and the second demand information is the actual demand information at the current moment. Interactive push data is generated based on the scene data at the current moment, the first demand information, and the second demand information, thereby improving the accuracy of data push.

[0101] Furthermore, based on any of the above embodiments, a sixth embodiment of the present invention for the intelligent voice interaction method for the elderly based on a large model is proposed, wherein the step of generating target voice data according to the interaction push data and the dialect database includes:

[0102] The interactive push data is standardized to obtain standard speech conversion data;

[0103] Prosody prediction is performed based on the standard speech conversion data to obtain speech feature data of the speech data. The speech feature data includes: the pause position, stress, intonation, and sentence duration of each sentence in the standard speech conversion data.

[0104] The target speech data is generated based on the speech feature data, the pronunciation dictionary, and the standard speech conversion data.

[0105] Specifically, the pronunciation dictionary needs to be updated in real time. In this embodiment, different speech generation methods can be executed for different interactive push data. Specifically, the aforementioned speech methods can use dialect speech synthesis technology, thereby solving the language barrier problem for elderly users. In addition to the above methods, the target speech data can also be obtained by recording relevant speech data.

[0106] Furthermore, this invention also proposes a large-model-based intelligent voice interaction device for the elderly, which includes:

[0107] The acquisition module is used to collect user voice data of elderly users and determine user text data based on the user voice data;

[0108] The identification module is used to determine the user profile of the elderly user based on the user text data and a preset large model;

[0109] The push module is used to determine interactive push data based on the user profile data;

[0110] The voice module is used to generate target voice data based on the interactive push data and the dialect database.

[0111] The large-model-based intelligent voice interaction device for the elderly can implement the steps of any of the above-described embodiments of the large-model-based intelligent voice interaction method for the elderly.

[0112] Furthermore, this invention also proposes an intelligent device, which includes: a memory, a processor, and a large-model-based intelligent voice interaction program for the elderly stored in the memory and executable on the processor. The large-model-based intelligent voice interaction program for the elderly is configured to implement the steps of any of the embodiments of the large-model-based intelligent voice interaction method for the elderly described above.

[0113] Furthermore, this embodiment of the invention also proposes a storage medium storing a large-model-based intelligent voice interaction program for the elderly. When the large-model-based intelligent voice interaction program for the elderly is executed by a processor, it implements the steps of any of the embodiments of the large-model-based intelligent voice interaction method for the elderly described above.

[0114] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0115] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0116] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0117] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A method for intelligent voice interaction for the elderly based on a large model, characterized in that, The intelligent voice interaction method for the elderly based on a large model includes the following steps: Collect user voice data from elderly users and determine user text data based on the user voice data; The user profile of the elderly user is determined based on the user text data and the preset large model; The interactive push data is determined based on the user profile data; Target speech data is generated based on the interactive push data and the dialect database.

2. The intelligent voice interaction method for the elderly based on a large model as described in claim 1, characterized in that, The number of user text data is multiple, and the step of determining the user profile of the elderly user based on the user text data and the preset large model includes: Based on the text time data and the first preset profile, multiple user text data are combined to obtain target text data, wherein the text time data is the time data corresponding to the user text data; First input data is generated based on the preset first portrait extraction instruction and the target text data; The first input data is input into the preset large model to obtain the model output, and the user profile is determined based on the model output.

3. The intelligent voice interaction method for the elderly based on a large model as described in claim 2, characterized in that, The text time data includes: text start time, text end time, and text duration. The step of combining multiple user text data based on the text time data and the first preset profile to obtain target text data includes: The text duration difference is obtained by comparing the time difference between two adjacent user text data and comparing the text duration of the two adjacent user text data. Based on the time difference and the text duration difference, determine the text time characteristics of two time-adjacent user text data; Based on the text time characteristics and the first preset profile, the multiple user text data are combined to obtain target text data; The time difference is the difference between the start time of the later user text data and the end time of the earlier user text data.

4. The intelligent voice interaction method for the elderly based on a large model as described in claim 3, characterized in that, The step of combining the multiple user text data according to the text time features and the first preset profile to obtain the target text data includes: When the text time feature is a first type of time feature, and one of the two user text data that are adjacent in time matches the portrait data of the first preset portrait, the user text data that does not match the portrait data of the first preset portrait in the two user text data that are adjacent in time is deleted. When the text time feature is a feature other than the first type of time feature, or when two user text data that are adjacent in time both match the portrait data of the first preset portrait, the two user text data that are adjacent in time are combined, and the combined text data is used as the updated user text data. When the number of user text data is 1, the user text data is used as the target text data; The first type of time feature is that the time difference is less than or equal to a preset time difference, and the text duration difference is less than or equal to a preset text duration difference.

5. The intelligent voice interaction method for the elderly based on a large model as described in claim 1, characterized in that, After the steps of inputting the first input data into the preset large model to obtain the model output and determining the user profile based on the model output, the method further includes: Calculate the cosine similarity between the user profile and the first preset profile; When the cosine similarity is less than or equal to the preset similarity, the first preset profile is updated according to the user profile; When the cosine similarity is greater than the preset similarity, the first preset portrait is updated according to the portrait database.

6. The intelligent voice interaction method for the elderly based on a large model as described in claim 1, characterized in that, The step of determining interactive push data based on the user profile data includes: The first demand information is determined based on the user profile data, and the second demand information is determined based on the user text data. The first demand information is emotional demand information, and the second demand information is the actual demand information at the current moment. Interactive push data is generated based on the current scene data, the primary demand information, and the secondary demand information.

7. The intelligent voice interaction method for the elderly based on a large model as described in claim 1, characterized in that, The step of generating target speech data based on the interactive push data and dialect database includes: The interactive push data is standardized to obtain standard speech conversion data; Prosody prediction is performed based on the standard speech conversion data to obtain speech feature data of the speech data. The speech feature data includes: the pause position, stress, intonation, and sentence duration of each sentence in the standard speech conversion data. The target speech data is generated based on the speech feature data, the pronunciation dictionary, and the standard speech conversion data.

8. A smart voice interaction device for the elderly based on a large model, characterized in that, The intelligent voice interaction device for the elderly based on a large model includes: The acquisition module is used to collect user voice data of elderly users and determine user text data based on the user voice data; The identification module is used to determine the user profile of the elderly user based on the user text data and a preset large model; The push module is used to determine interactive push data based on the user profile data; The voice module is used to generate target voice data based on the interactive push data and the dialect database.

9. A smart device, characterized in that, The intelligent device includes: a memory, a processor, and a large-model-based intelligent voice interaction program for the elderly stored in the memory and executable on the processor, wherein the large-model-based intelligent voice interaction program for the elderly is configured to implement the steps of the large-model-based intelligent voice interaction method for the elderly as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a large-model-based intelligent voice interaction program for the elderly. When the large-model-based intelligent voice interaction program for the elderly is executed by the processor, it implements the steps of the large-model-based intelligent voice interaction method for the elderly as described in any one of claims 1 to 7.