System and method for language learning resource enhancement

WO2026183177A1PCT designated stage Publication Date: 2026-09-03HOWARD UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/016586
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-26
Filing Date
2026-02-25
Publication Date
2026-09-03

Smart Images

  • Figure US2026016586_03092026_PF_FP_ABST
    Figure US2026016586_03092026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are systems, methods, and devices for enhancing language learning resources. According to example embodiments, the system may include: a memory storage storing computer-executable instructions; and at least one processor communicatively coupled to the memory storage, wherein the at least one processor may be configured to execute the instructions to: obtain a plurality of recorded voice lines of a secondary user; train an artificial intelligence (AI) voice assistant based on the obtained plurality of voice lines of the secondary user; and provide a generated voice line to a primary user using the AI voice assistant, such that the generated voice line is provided with a voice of the secondary user; wherein the generated voice line may be associated with language learning process.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR LANGUAGE LEARNING RESOURCE ENHANCEMENTTECHNICAL FIELD

[0001] Example embodiments of the present disclosure relate to language learning, and more specifically, relate to enhancement in resources for language learning.BACKGROUND

[0002] Language learning resources, such as language learning applications and software, have become increasingly relevant as the world becomes increasingly diverse and people from different cultures and ethnicities begin to merge and integrate with each other in the modern age.

[0003] One implementation of language learning applications that has been gathering attention in recent history involves the use of an artificial intelligence (Al) voice assistant. Such Al voice assistant implemented in language learning applications may be trained (e.g., using various machine learning (ML) techniques) to receive voice commands from a user in a natural language, process the received voice commands, and provide appropriate responses to the user in order to assist and facilitate learning experience.

[0004] The use of Al voice assistants implemented in language learning applications has been particularly advantageous in helping children during their early stages in language development, as well as children with speech challenges and communication disorders.SUMMARY

[0005] Example embodiments consistent with the present disclosure allow the children to receive and hear voice lines speaking and pronouncing certain words with the voices of a family member, which improves the language learning experience and effectiveness for the children.

[0006] According to example embodiments, a system is provided. The system may include: a memory storage storing computer-executable instructions; and at least one processor communicatively coupled to the memory storage, wherein the at least one processor may be configured to execute the instructions to: obtain a plurality of recorded voice lines of a secondary user; train an artificial intelligence (Al) voice assistant based on the obtained plurality of voice lines of the secondary user; and provide a generated voice line to a primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the secondary user; wherein the generated voice line may be associated with language learning process.

[0007] According to example embodiments, the second user may include a family member of the primary user.

[0008] According to example embodiments, the at least one processor may be configured to execute the instructions to obtain the plurality of recorded voice lines of the secondary user by: displaying a plurality of training images associated with a plurality of words to the secondary user; receiving a plurality of voice lines corresponding to the plurality of words associated with the plurality of training images from the secondary user; and storing the plurality of voice lines received from the secondary user.

[0009] According to example embodiments, the plurality of training images may include at least one predefined training image and at least one training image received from the secondary user.

[0010] According to example embodiments, the at least one processor may be configured to execute the instructions to train the Al voice assistant based on the obtained plurality of voice lines of the secondary user by: training a machine learning (ML) model based on the obtained plurality of voice lines of the secondary user; and configuring the Al voice assistant based on the trained ML model.

[0011] According to example embodiments, the at least one processor may be configured to execute the instructions to provide the generated voice line to the primary user using the Al voice assistant by; displaying a learning image associated with a word to the primary user; generating a learning voice line corresponding to the word associated with the learning image using the Al voice assistant; and providing the generated learning voice line to the primary user, such that the generated learning voice line is provided with the voice of the secondary user.

[0012] According to example embodiments, the at least one processor may be further configured to execute the instructions to provide the generated voice line to the primary user using the Al voice assistant by: in response to providing the generated learning voice line to the primary user, receiving an input voice line from the primary user; determining whether the input voice line corresponds to the word associated with the learning image; in response to determining that the input voice line corresponds to the word associated with the learning image, generating a positive affirmation voice line using the Al voice assistant, and providing the generated positive affirmation voice line to the primary user, such that the generated positive affirmation voice line is provided with the voice of the secondary user; and in response to determining that the input voice line does not correspond to the word associated with the learning image, providing the generated learning voice line to the primary user, such that the generated learning voice line is provided with the voiceof the secondary user.

[0013] According to example embodiments, the at least one processor may be configured to execute the instructions to provide the generated voice line to the primary user using the Al voice assistant by: displaying a plurality of secondary users who have previously provided the plurality of recorded voice lines; receiving a selection input selecting one of the displayed plurality of secondary users; and providing the generated voice line to the primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the selected one of the displayed plurality of secondary users.

[0014] According to example embodiments, a method is provided. The method may include: obtaining a plurality of recorded voice lines of a secondary user; training an artificial intelligence (Al) voice assistant based on the obtained plurality of voice lines of the secondary user; and providing a generated voice line to a primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the secondary user; wherein the generated voice line may be associated with language learning process.

[0015] According to example embodiments, the second user may include a family member of the primary user.

[0016] According to example embodiments, the obtaining the plurality of recorded voice lines of the secondary user may include: displaying a plurality of training images associated with a plurality of words to the secondary user; receiving a plurality of voice lines corresponding to the plurality of words associated with the plurality of training images from the secondary user; and storing the plurality of voice lines received from the secondary user.

[0017] According to example embodiments, the plurality of training images may includeat least one predefined training image and at least one training image received from the secondary user.

[0018] According to example embodiments, the training the Al voice assistant based on the obtained plurality of voice lines of the secondary user may include: training a machine learning (ML) model based on the obtained plurality of voice lines of the secondary user; and configuring the Al voice assistant based on the trained ML model.

[0019] According to example embodiments, the providing the generated voice line to the primary user using the Al voice assistant may include: displaying a learning image associated with a word to the primary user; generating a learning voice line corresponding to the word associated with the learning image using the Al voice assistant; and providing the generated learning voice line to the primary user, such that the generated learning voice line is provided with the voice of the secondary user.

[0020] According to example embodiments, the providing the generated voice line to the primary user using the Al voice assistant may further include: in response to providing the generated learning voice line to the primary user, receiving an input voice line from the primary user; determining whether the input voice line corresponds to the word associated with the learning image; in response to determining that the input voice line corresponds to the word associated with the learning image, generating a positive affirmation voice line using the Al voice assistant, and providing the generated positive affirmation voice line to the primary user, such that the generated positive affirmation voice line is provided with the voice of the secondary user; and in response to determining that the input voice line does not correspond to the word associated with the learning image, providing the generated learning voice line to the primary user, such that the generatedlearning voice line is provided with the voice of the secondary user.

[0021] According to example embodiments, the providing the generated voice line to the primary user using the Al voice assistant may include: displaying a plurality of secondary users who have previously provided the plurality of recorded voice lines; receiving a selection input selecting one of the displayed plurality of secondary users; and providing the generated voice line to the primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the selected one of the displayed plurality of secondary users.

[0022] According to example embodiments, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium may have recorded thereon instructions executable by at least one processor to cause the at least one processor to perform a method including: obtaining a plurality of recorded voice lines of a secondary user; training an artificial intelligence (Al) voice assistant based on the obtained plurality of voice lines of the secondary user; and providing a generated voice line to a primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the secondary user; wherein the generated voice line may be associated with language learning process.

[0023] According to example embodiments, the second user may include a family member of the primary user.

[0024] According to example embodiments, the obtaining the plurality of recorded voice lines of the secondary user may include: displaying a plurality of training images associated with a plurality of words to the secondary user; receiving a plurality of voice lines corresponding to the plurality of words associated with the plurality of training images from the secondary user; and storing the plurality of voice lines received from the secondary user.

[0025] According to example embodiments, the plurality of training images may include at least one predefined training image and at least one training image received from the secondary user.

[0026] Additional aspects will be set forth in part in the description that follows and, in part, will be apparent from the description, or may be realized by practice of the presented embodiments of the disclosure.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Features, advantages, and significance of exemplary embodiments of the disclosure will be described below with reference to the accompanying drawings, in which like reference numerals denote like elements, and wherein:

[0028] FIG. 1 illustrates a block diagram of an example system configuration for enhancing language learning resource, according to one or more example embodiments;

[0029] FIG. 2 illustrates a flow diagram of an example method for enhancing language learning resource, according to one or more example embodiments;

[0030] FIG. 3 illustrates a flow diagram of an example method for obtaining a plurality of recorded voice lines of a secondary user, according to one or more example embodiments;

[0031] FIG. 4A and FIG. 4B illustrate examples of graphical user interface (GUI) for obtaining a plurality of recorded voice lines of a secondary user, according to one or more example embodiments;

[0032] FIG. 5 illustrates a flow diagram of an example method for training an artificial intelligence (Al) voice assistant, according to one or more example embodiments;

[0033] FIG. 6 illustrates a flow diagram of an example method for providing a generated voice line to a primary user, according to one or more example embodiments;

[0034] FIG. 7 illustrates an example graphical user interface (GUI) for providing a generated voice line to a primary user, according to one or more example embodiments;

[0035] FIG. 8 illustrates a flow diagram of an example method for providing a generated voice line to a primary user, according to one or more example embodiments; and

[0036] FIG. 9 illustrates a block diagram of example components in a system, according to one or more example embodiments.DETAILED DESCRIPTION

[0037] The following detailed description of exemplary embodiments refers to the accompanying drawings. The foregoing disclosure provides illustration and description but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementations. Further, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, in the flowcharts and descriptions of operations provided below, it is understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least in part), and the order of one or more operations may be switched.

[0038] Even though particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure ofpossible implementations. Tn fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of possible implementations includes each dependent claim in combination with every other claim in the claim set.

[0039] No element, act, or instruction used herein should be construed as critical or essential unless explicitly described as such. Also, as used herein, the articles “a” and “an” are intended to include one or more items, and may be used interchangeably with “one or more.” Where only one item is intended, the term “one” or similar language is used. Also, as used herein, the terms “has,” “have,” “having,” “include,” “including,” or the like are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “based, at least in part, on” unless explicitly stated otherwise. Furthermore, expressions such as “[A] and / or [B]”, “at least one of [A] and [B]” or “at least one of [A] or [B]” are to be understood as including only A, only B, or both A and B.

[0040] Expressions such as “at least one processor,” where configured to implement a plurality of operations, execute a plurality of instructions, etc., are to be understood as a single processor implementing the plurality of operations, etc., or certain of plural processors implementing certain (but not necessarily all) of the plurality of operations, etc.

[0041] Reference throughout this specification to “one embodiment,” “an embodiment,” “non-limiting exemplary embodiment,” or similar language means that a particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of the present solution. Thus, the phrases “in one embodiment”, “in anembodiment,” “in one non-limiting exemplary embodiment,” and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

[0042] Further, the described features, advantages, and characteristics of the present disclosure may be combined in any suitable manner in one or more example embodiments. One skilled in the relevant art will recognize, in light of the description herein, that the present disclosure can be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments of the present disclosure.

[0043] As explained above, the use of an artificial intelligence (Al) voice assistant in language learning applications has been gathering attention in recent history.

[0044] An Al voice assistant can be trained using pre-recorded voices during development, which results in the voices of many Al voice assistants being racially biased to white / western cultures, accents, and dialogues.

[0045] Such racial bias of Al voice assistants may have negative effect on the learning ability and experience of children from a different ethnicity during their language acquisition process, as children can have significant improvements in engagement with Al voice assistant from the same ethnicity in comparison to Al voice assistant from different ethnicity. The above issue may be compounded even further for children with speech challenges and communication disorders.

[0046] Similar issues have been studied where significant shortages of speech-language pathologists (SLP) result in many SLPs being overwhelmingly white / western. The SLP’s lack of cultural divergence may have led to many children from different ethnicities, such as AfricanAmericans, Indians, Asians, and the like, being denied services or being misdiagnosed.

[0047] Accordingly, there is a need for a method for providing language learning resources, such as Al voice assistants, that are able to adapt and cater to children from different ethnicities.

[0048] It is contemplated that features, advantages, and significances of example embodiments described herein are merely a portion of the present disclosure, and are not intended to be exhaustive or to limit the scope of the present disclosure. Further descriptions of the features, components, configuration, operations, and implementations of the example embodiments of the present disclosure are provided in the following.

[0049] FIG. 1 illustrates a block diagram of an example system configuration 100 for enhancing language learning resource, according to one or more example embodiments. As illustrated in FIG. 1, system configuration 100 may include a language learning (LL) system 110, a primary user 120, and a secondary user 130, although it is contemplated that the system configuration 100 may include more or less components than illustrated in FIG. 1, without departing from the scope of the present disclosure.

[0050] LL system 110 may include an apparatus, a system, a platform, a module, or the like, which may be configured to perform one or more operations or actions for enhancing language learning resource. According to example embodiments, the LL system 110 may include and / or be implemented on a user equipment, such as a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smart phone, a radiotelephone, etc.), a wearable device (e.g., a pair of smart glasses or a smart watch), a SIM-based device, and / or a similar device.

[0051] According to example embodiments, the LL system 110 may also include and / or be associated with a language learning application. The language learning application may refer to an application (software, platform, etc.) configured to assist a user in learning a language.

[0052] Example operations performable by the LL system 110 for enhancing language learning resource are described below with reference to FIG. 2 to FIG. 8. Further, several example components which may be included in the LL system 110, according to one or more example embodiments, are described below with reference to FIG. 9.

[0053] Primary user 120 may include a primary (main) user of the LL system 110 and / or the language learning application. In particular, the primary user 120 may include a user who is utilizing the LL system 110 and / or the language learning application to learn a language. According to example embodiments, the primary user 120 may include children. In particular, for example, the primary user 120 may include children during early stages of speech development (e.g., around age 3-5).

[0054] Secondary user 130 may include a secondary user of the LL system 110 and / or the language learning application. In particular, the secondary user 130 may include a user who is utilizing the LL system 110 and / or the language learning application to assist the primary user 120 to learn a language. According to example embodiments, the secondary user 130 may include a family member of the primary user 120. For example, the secondary user 130 may be the father / mother, while the primary user 120 may be the son / daughter.

[0055] In this regard, according to example embodiments, the primary and secondary users 120,130 may be communicatively coupled to the LL system 110. In particular, the primary and secondary users 120,130 may provide inputs, such as voice lines, text, and the like, to the LLsystem 110 via input components of the LL system 110 (e.g., microphone, keyboard, touch screen, etc.). Similarly, the LL system 110 may provide outputs, such as voice lines, text, images, and the like, to the primary and secondary users 120,130 via output components of the LL system 110 (e g., speakers, display screens, etc.)

[0056] In the following, several example operations performable by the LL system of the present disclosure are described with reference to FIG. 2 to FIG. 8.

[0057] FIG. 2 illustrates a flow diagram of an example method 200 for enhancing language learning resource, according to one or more example embodiments. One or more operations in method 200 may be performed by at least one processor (e.g., processor 912) of the LL system.

[0058] As illustrated in FIG. 2, at operation S210, the at least one processor may be instructed by program code to obtain a plurality of recorded voice lines of a secondary user.

[0059] According to example embodiments, the at least one processor may be instructed by program code to obtain a plurality of recorded voice lines of the secondary user by: displaying a plurality of training images associated with a plurality of words to the secondary user; receiving a plurality of voice lines corresponding to the plurality of words associated with the plurality of training images from the secondary user; and storing the plurality of voice lines of the secondary user. According to example embodiments, the at least one processor may be instructed by program code to obtain the plurality of recorded voice lines of the secondary user via a graphical user interface (GUI).

[0060] It is understood that the term “image” used herein may include visual images as well as texts, while the term “word” used herein may include words as well as numbers, letters, and the like, associated with language learning unless explicitly stated otherwise. Further, the term“voice” used herein may include voices, as well as pitch, intonation, timbre, phoneme length, and the like associated with speech.

[0061] Examples of operations for obtaining the plurality of recorded voice lines of the secondary user are described below with reference to FIG. 3. The method then proceeds to operation S220.

[0062] At operation S220, the at least one processor may be instructed by program code to train an artificial intelligence (Al) voice assistant based on the obtained plurality of voice lines of the secondary user.

[0063] According to example embodiments, the at least one processor is instructed by program code to train the Al voice assistant based on the obtained plurality of voice lines of the secondary user by: training a machine learning (ML) model based on the obtained plurality of voice lines of the secondary user; and configuring the Al voice assistant based on the trained ML model.

[0064] Examples of operations for training the Al voice assistant are described below with reference to FIG. 5. The method then proceeds to operation S230.

[0065] At operation S230, the at least one processor may be instructed by program code to provide a generated voice line to a primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the secondary user. The generated voice line may be associated with language learning process. For example, the generated voice line may correspond to a word, a sentence, and the like which the primary user would like to practice speaking as part of a language learning process. Further, the secondary user may include a family member of the primary user.

[0066] According to example embodiments, the at least one processor may be instructed by program code to provide the generated voice line to the primary user using the Al voice assistant by: displaying a learning image associated with a word to the primary user; generating a learning voice line corresponding to the word associated with the learning image using the Al voice assistant; and providing the generated learning voice line to the primary user, such that the generated learning voice line is provided with the voice of the secondary user.

[0067] Subsequently, the at least one processor may be further instructed by program code to provide the generated voice line to the primary user using the Al voice assistant by: receiving an input voice line from the primary user; and determining whether the input voice line corresponds to the word associated with the learning image. In this regard, in response to determining that the input voice line corresponds to the word associated with the learning image, the at least one processor may be instructed by program code to generate a positive affirmation voice line using the Al voice assistant, and provide the generated positive affirmation voice line to the primary user, such that the generated positive affirmation voice line is provided with the voice of the secondary user. On the other hand, in response to determining that the input voice line does not correspond to the word associated with the learning image, the at least one processor may be instructed by program code to provide (re-provide) the generated learning voice line to the primary user, such that the generated learning voice line is provided with the voice of the secondary user.

[0068] According to example embodiments, the at least one processor may be instructed by program code to provide the generated voice line to the primary user using the Al voice assistant by: displaying a learning sentence to the primary user; generating a learning voice linecorresponding to the learning sentence using the Al voice assistant; and providing the generated learning voice line to the primary user, such that the generated learning voice line is provided with the voice of the secondary user.

[0069] According to example embodiments, the at least one processor may be instructed by program code to provide the generated voice line to the primary user using the Al voice assistant by: displaying a plurality of secondary users who have previously provided the plurality of recorded voice lines; receiving a selection input selecting one of the displayed plurality of secondary users; and provide the generated voice line to the primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the selected one of the displayed plurality of secondary users.

[0070] Examples of operations for providing the generated voice line to the primary user are described below with reference to FIG. 6 and FIG. 8.

[0071] Upon performing operation S230, the method 200 may be ended or be terminated. Alternatively, method 200 may return to operation S210, such that the at least one processor may be instructed by program code to repeatedly perform, for at least a predetermined amount of time, the obtaining the plurality of recorded voice lines (at operation S210), the training the Al voice assistant (at operation S220), and the providing the generated voice line (at operation S230).

[0072] For instance, the at least one processor may continuously (or periodically) receive a plurality of recorded voice lines from different family members (plurality of secondary users) of the primary user, and then restart the obtaining the plurality of recorded voice lines (at operation S210), the training the Al voice assistant (at operation S220), and the providing the generated voice line (at operation S230).

[0073] Accordingly, the above process for enhancing language learning resource allows the children (primary user) to receive and hear voice lines speaking and pronouncing certain words with the voices of the family member (secondary user), which improves language learning experience and effectiveness for the children.

[0074] In particular, the above process focuses on facilitating early language development, cultural inclusivity, and personalized learning experiences. More specifically, the utilization of a family member’s voice enhances the personalized language learning experience for the children, which ultimately provides inclusive options that cater to the needs of individual children from all communities, while also allowing family members to participate in their children’s language learning journey thereby fostering stronger connection for the children with their culture. The above process also facilitates and improves the children’s ability to learn multiple languages.

[0075] FIG. 3 illustrates a flow diagram of an example method 300 for obtaining a plurality of recorded voice lines of a secondary user, according to one or more example embodiments. One or more operations of method 300 may be part of operation S210 in method 200, and may be performed by at least one processor of the LL system. Further, one or more operations of method 300 may be performed via a graphical user interface (GUI).

[0076] As illustrated in FIG. 3, at operation S310, the at least one processor may be instructed by program code to display a plurality of training images associated with a plurality of words to the secondary user. The plurality of training images may refer to a set of images used to obtain voice lines from the secondary users (see below in operations S320 and S33O).

[0077] For example, the plurality of training images may include an image of a dog (or an image of a text “Dog”), where such image may be associated with the word “dog”. In anotherexample, the plurality of training images may include an image of a number “1”, where such image may be associated with the word “one”. The method then proceeds to operation S320.

[0078] At operation S320, the at least one processor may be instructed by program code to receive a plurality of voice lines corresponding to the plurality of words associated with the plurality of training images from the secondary user.

[0079] In particular, according to example embodiments, the at least one processor may be instructed by program code to display the plurality of training images associated with the plurality of words to the secondary user, receive a selection input (e.g., from the secondary user) selecting one of the plurality of training images, and receive a voice line corresponding to the word associated with the selected one of the plurality of training images from the secondary user. In addition, the at least one processor may be instructed by program code to display the selected one of the plurality of training images to the secondary user.

[0080] For example, the at least one processor may receive a selection input from the secondary user selecting an image of a dog (or an image of a text “Dog”), and then receive a voice line saying the word “dog” (i.e., voice line corresponding to the word associated with the selected one training image) from the secondary user.

[0081] Alternatively or in addition to the above, the at least one processor may be instructed by program code to randomly select one of the plurality of training images, display the selected one of the plurality of training images to the secondary user, and receive a voice line corresponding to the word associated with the selected one of the plurality of training images from the secondary user.

[0082] The above processes may be repeated until a plurality of voice lines correspondingto the plurality of words associated with the plurality of training images are received from the secondary user.

[0083] According to example embodiments, the plurality of training images may be predefined (i.e., predefined training image). According to example embodiments, the plurality of training images may be received from the secondary user and / or the primary user. According to example embodiments, the plurality of training images may include at least one predefined training image and at least one training image received from the secondary user and / or the primary user.

[0084] For example, the plurality of words 310 shown in FIG. 4A may be predefined, and the at least one processor may be instructed by program code to receive a new word (i.e., new training image) “Bat” from the secondary user (e.g., via an input interface element displayed on the GUI for receiving images, texts, etc.), and add said new word to the plurality of words 310 (i.e., a pool of existing training images). Accordingly, said new word can be selected and voice line for said new word can be received in the similar manner as described above in relation to operation S310 and S320 as well as FIG. 4A and FIG. 4B. The method then proceeds to operation S330.

[0085] At operation S330, the at least one processor may be instructed by program code to store the plurality of voice lines received from the secondary user.

[0086] FIG. 4A and FIG. 4B illustrate examples of graphical user interface (GUI) for obtaining a plurality of recorded voice lines of a secondary user, according to one or more example embodiments.

[0087] As shown in FIG. 4A, a graphical user interface (GUI) 400 may be generated by the at least one processor and displayed to the secondary user in the similar manner as describedabove in relation to operations S310 and S320 in method 300, where the GUT 400 may include a plurality of words 410 (i.e., plurality of training images). In the example shown in FIG. 4A, the plurality of words 410 includes “Bug”, “Blog”, “Bold”, “Browser”, and “Battery” (i.e., plurality of words associated with the plurality of training images). Here, in response to receiving a selection input from the secondary user selecting the word “Blog” from the plurality of words 410, the GUI 400 may be updated by the at least one processor to display the selected word 420 “Blog” as well as a recording button 430 as shown in FIG. 4B.

[0088] In this regard, in response to the secondary user pressing on the recording button 430 and speaking (i.e., providing a voice line) the word “Blog”, the at least one processor may receive the voice line of the secondary user speaking the word “Blog” (i.e., voice line corresponding to the word associated with the selected one training image), in the similar manner as described above in relation to operation S320.

[0089] Subsequently, the at least one processor may store the received voice line, in the similar manner as described above in relation to operations S330.

[0090] It is understood that the configurations illustrated in FIG. 4A and FIG. 4B are simplified for descriptive purpose, and is not intended to limit the scope of the present disclosure in any way.

[0091] For example, the graphical user interface may include additional interface elements other than the recording button 430, such as a playback button that allows the secondary user to play the recorded voice line, a save button that allows the secondary to save (store) the recorded voice line, a delete button that allows the secondary user to delete the recorded voice line, and the like. Further, the plurality of words 310 and the selected word 320 may be shown on the GUI 300a visual images instead of texts. Furthermore, the graphical user interface may group the plurality of words 310 into categories (e.g., type, alphabetical, etc.), where the graphical user interface may first display the categories to the secondary user, receive a selection input from the secondary user selecting one of the categories, and then display the plurality of words 310 within the selected category (e.g., FIG. 4A may show a plurality of words 310 within the category of alphabetical “B”). Further still, the voice line may be received from the secondary user automatically instead of in response to the secondary user pressing on the record button 430.

[0092] FIG. 5 illustrates a flow diagram of an example method 500 for training an artificial intelligence (Al) voice assistant, according to one or more example embodiments. One or more operations of method 500 may be part of operation S220 in method 200, and may be performed by at least one processor of the LL system.

[0093] As illustrated in FIG. 5, at operation S510, the at least one processor may be instructed by program code to train a machine learning (ML) model based on the obtained plurality of voice lines of the secondary user.

[0094] In particular, according to example embodiments, the ML model may be trained to analyze and extract unique features from the obtained plurality of voice lines of the secondary user, such as pitch, intonation, timbre, phoneme length, and the like from the obtained plurality of voice lines of the secondary user. The extracted unique features may together form a representation of a unique voice of the secondary user, such that the ML model may be utilized to reproduce the unique voice of the secondary user.

[0095] In this regard, according to example embodiments, the at least one processor may be instructed by program code to train a plurality of ML models based on obtained plurality ofvoice lines of a plurality of secondary users, where different ML models correspond to the voice of different secondary users. For example, the at least one processor may be instructed by program code to train a first ML model based on obtained plurality of voice lines of a first secondary user (e g., father), and train a second ML model based on obtained plurality of voice lines of a second secondary user (e.g., mother), such that the first ML model may be utilized to reproduce the unique voice of the first secondary user and the second ML model may be utilized to reproduce the unique voice of the second secondary user.

[0096] In this regard, the ML model may include any kind of ML model configured to analyze, extract, and mimic unique features of human voices. Further, it is understood that the ML model may be trained using any appropriate method, including, but not limited to, deep learning algorithms, deep neural network (DNN), recurrent neural network (RNN), convolutional neural network (CNN), and the like. The method then proceeds to operation S520.

[0097] At operation S520, the at least one processor may be instructed by program code to configure the Al voice assistant based on the trained ML model.

[0098] According to example embodiments, the Al voice assistant may include an application, a software, a program, and the like that utilizes artificial intelligence (Al) in order to understand and perform operations / tasks in response to commands from a user. Such Al voice assistant may utilize various techniques in order to understand and perform operations / tasks in response to commands from a user, including, but not limited to, natural language processing (NLP), speech recognition, and the like.

[0099] In this regard, the Al voice assistant may be configured to provide a generated voice line (see description below in relation to FIG. 6 and FIG. 7). As such, the Al voice assistant maybe configured based on the trained ML model (which represents a unique voice of the secondary user) to produce a generated voice line with the unique voice of the secondary user.

[0100] FIG. 6 illustrates a flow diagram of an example method 600 for providing a generated voice line to a primary user, according to one or more example embodiments. One or more operations of method 600 may be part of operation S230 in method 200, and may be performed by at least one processor of the LL system. Further, one or more operations of method 600 may be performed via a graphical user interface (GUI).

[0101] As illustrated in FIG. 6, at operation 610, the at least one processor may be instructed by program code to display a learning image associated with a word to the primary user. The learning image may refer to an image used to assist the primary user in learning a language.

[0102] For example, the learning image may include an image of a dog (or an image of a text “Dog”), where such image may be associated with the word “dog”. In another example, the learning image may include an image of a number “1”, where such image may be associated with the word “one”.

[0103] According to example embodiments, the at least one processor may be instructed by program code to display the learning image associated with the word to the primary user by: displaying a plurality of learning images associated with a plurality of words to the primary user, receiving a selection input (e.g., from the primary user or from the secondary user) selecting one of the plurality of learning images, and displaying the selected one of the plurality of learning images to the primary user.

[0104] According to example embodiments, the at least one processor may be instructed by program code to display the learning image associated with the word to the primary user by:randomly selecting one of a plurality of learning images associated with a plurality of words, and displaying the selected one of the plurality of learning images to the primary user.

[0105] In this regard, according to example embodiments, the plurality of learning images may be predefined (i.e., predefined learning image). According to example embodiments, the plurality of learning images may be received from the secondary user and / or the primary user. According to example embodiments, the plurality of learning images may be automatically generated using the Al voice assistant. According to example embodiments, the plurality of learning images may include at least two of predefined learning image, learning image received from the secondary user and / or the primary user, and learning image automatically generated using the Al voice assistant. The method then proceeds to operation S620.

[0106] At operation 620, the at least one processor may be instructed by program code to generate a learning voice line corresponding to the word associated with the learning image using the Al voice assistant.

[0107] In this regard, since the Al voice assistant is trained based on the obtained plurality of voice lines of the secondary user, as described above in relation to method 500, the learning voice line generated using such Al voice assistant may have a voice of the secondary user. The method then proceeds to operation S630

[0108] At operation 630, the at least one processor may be instructed by program code to provide the generated learning voice line to the primary user. In this regard, since the generated learning voice line may have the voice of the secondary user, the generated learning voice line may be provided with the voice of the secondary user.

[0109] For example, if the learning image includes an image of a dog (or an image of atext “Dog”) which is associated with the word “dog”, the at least one processor may be instructed by program code to generate and provide a learning voice line saying the word “dog” with the voice of the secondary user to the primary user.

[0110] In this regard, the above processes may allow the children (primary user) to receive and hear voice lines speaking and pronouncing certain words with the voices of the family member (secondary user), which may improve language learning experience and effectiveness for the children. The method then proceeds to operation S640.

[0111] At operation 640, the at least one processor may be instructed by program code to receive an input voice line from the primary user. The input voice line may refer to a voice line of the primary user attempting to speak the word associated with the learning image.

[0112] In particular, in response to hearing the generated learning voice line corresponding to the word associated with the learning image (which is also displayed to the primary user), the primary user may attempt to mimic and learn the pronunciation of such word.

[0113] Such process encourages and enables the primary user (children) to compare their own voices with those of their family member, thereby improve language learning experience and effectiveness for the children while also fostering a strong sense of connection to their cultural identity. The method then proceeds to operation S650.

[0114] At operation 650, the at least one processor may be instructed by program code to determine whether the input voice line corresponds to the word associated with the learning image.

[0115] According to example embodiments, the at least one processor may be instructed by program code to determine whether the input voice line corresponds to the word associated with the learning image by determining whether a word associated with the input voice linematches with the word associated with the learning image.

[0116] According to example embodiments, the at least one processor may be instructed by program code to determine whether the input voice line corresponds to the word associated with the learning image by determining whether pronunciations (e.g., pitch, intonation, timbre, phoneme length, and the like) of a word associated with the input voice line matches with pronunciations of the word associated with the learning image (i.e., the pronunciations of the generated learning voice line).

[0117] According to example embodiments, the at least one processor may determine whether the input voice line corresponds to the word associated with the learning image using the Al voice assistant.

[0118] It is understood that the at least one processor may utilize any appropriate methods in order to determine whether the input voice line corresponds to the word associated with the learning image. For example, the at least one processor may utilize speech recognition to convert the received voice line from an acoustic form into a text form, natural language processing (NLP) to understand the received voice line, and the like.

[0119] In this regard, in response to determining that the input voice line corresponds to the word associated with the learning image, the at least one processor may determine that the primary user has correctly learned the word associated with the learning image, and the method proceeds to operation S660. On the other hand, in response to determining that the input voice line does not correspond to the word associated with the learning image, the at least one processor may determine that the primary user has not yet correctly learn the word associated with the learning image, and the method returns to operation S630 to re-provide the generated learning voice line tothe primary user.

[0120] At operation 660, the at least one processor may be instructed by program code to generate a positive affirmation voice line using the Al voice assistant. The positive affirmation voice line may be generated using the Al voice assistant in the similar manner as the learning voice line. The method then proceeds to operation S670.

[0121] At operation S670, the at least one processor may be instructed by program code to provide the generated positive affirmation voice line to the primary user. In this regard, since the generated positive affirmation voice line may have the voice of the secondary user, the generated positive affirmation voice line may be provided with the voice of the secondary user.

[0122] In this regard, the above processes may allow the children (primary user) to receive positive affirmation in the voice of the family member (secondary user), which may provide and reinforce positive feedback and thereby further improving language learning experience and effectiveness for the children.

[0123] In addition, the operations in method 600 may be extended to encompass a combination of words, such as sentences, paragraphs, articles, books, and the like, such that the at least one processor may generate and provide learning voice line corresponding to the above using the Al voice assistant.

[0124] FIG. 7 illustrates an example graphical user interface (GUI) for providing a generated voice line to a primary user, according to one or more example embodiments. As shown in FIG. 7, a graphical user interface (GUI) 700 may be generated by the at least one processor and displayed to the primary user in the similar manner as described above in relation to operation S610 in method 600, where the GUI 700 may include a word 710 ’’Bill” (i.e., learning image), aplay-back button 720, and a record button 730.

[0125] Here, in response to the primary user pressing on the play-back button 720, the at least one processor may generate and provide the learning voice line to the primary user, in the similar manner as described above in relation to operations S620 and S630. For example, the at least one processor may generate and provide the learning voice line saying “Bill” with the voice of the secondary user.

[0126] Further, in response to the primary user pressing on the record button 730, the at least one processor may receive the input voice line from the primary user, in the similar manner as described above in relation to operation S640. For example, the at least one processor may receive the input voice line from the primary user attempting to say “Bill”.

[0127] Subsequently, the at least one processor may determine whether the input voice line corresponds to the word associated with the learning image, and then provide the positive affirmation voice line to the primary user or re-provide the learning voice line to the primary user, in the similar manner as described above in relation to operations S650 to S670.

[0128] It is understood that the configurations illustrated in FIG. 7 is simplified for descriptive purpose, and is not intended to limit the scope of the present disclosure in any way. For example, the graphical user interface may include additional interface elements other than the play-back button 720 and the record button 730, the word 710 may be shown on the GUI 700 a visual images instead of texts, the input voice line may be received from the primary user automatically instead of in response to the primary user to pressing on the record button 730, and the like.

[0129] FIG. 8 illustrates a flow diagram of an example method 800 for providing agenerated voice line to a primary user, according to one or more example embodiments. One or more operations of method 800 may be part of operation S230 in method 200, and may be performed by at least one processor of the LL system. Further, one or more operations of method 800 may be performed via a graphical user interface (GUI).

[0130] As illustrated in FIG. 8, at operation 810, the at least one processor may be instructed by program code to display a plurality of secondary users who have previously provided the plurality of recorded voice lines.

[0131] In particular, according to example embodiments, operations S220 and S230 described above in relation to method 200 as well as operations described above in relation to method 300 and 500 may be performed with a plurality of secondary users, such that the Al voice assistant may generate voice lines with the voices of the plurality of secondary users.

[0132] Accordingly, the at least one processor may display such plurality of secondary users who have previously provided the plurality of recorded voice lines (and whom the Al voice assistant may generate voice lines with their voices). In this regard, the plurality of secondary users may be displayed as icons, pictures, names, roles (father, mother, etc.), and the like. The method then proceeds to operation S820.

[0133] At operation 820, the at least one processor may be instructed by program code to receive a selection input selecting one of the displayed plurality of secondary users. The selection input may be received from the primary user or the secondary user. The method then proceeds to operation S830.

[0134] At operation 830, the at least one processor may be instructed by program code to provide the generated voice line (e.g., learning voice line) to the primary user using the Al voiceassistant. Tn this regard, the generated voice line may have the voice of the selected one of the displayed plurality of secondary users, such that the generated voice line may be provided with the voice of selected one of the displayed plurality of secondary users.

[0135] For example, the at least one processor may display icons of father and mother (plurality of secondary users) during operation S810. Then, the primary user may provide a selection input selecting the mother during operation S820. In response, the at least one processor may generate and provide a voice line to the primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the mother (selected one of the displayed plurality of secondary users).

[0136] In this regard, the above processes may allow the children (primary user) to selectively receive and hear voice lines speaking and pronouncing certain words with the voices of a particular family member (secondary user), which may further improve language learning experience and effectiveness for the children.

[0137] FIG. 9 illustrates a block diagram of example components in a system 900, according to one or more example embodiments. The system 900 may correspond to the LL system 110 in FIG. 1, thus the features associated with the LL system 110 and the system 900 may be similarly applicable to each other, unless being explicitly described otherwise.

[0138] As illustrated in FIG. 9, the system 910 may include at least one bus 911, at least one processor 912, at least one memory 913, at least one storage component 914, at least one input component 915, at least one output component 916, and at least one communication interface 917.

[0139] It is contemplated that the system 910 may include more or less components than illustrated in FIG. 9, without departing from the scope of the present disclosure. For instance, insome embodiments, the system 910 may include a plurality of storage components 914, the input component 915 and the output component 916 may be implemented as a transceiver component, the memory 913 and storage component 914 may be implemented as a memory storage, and the like.

[0140] The bus 911 may be configured to facilitate or enable communications among the components of the system 910. Specifically, the bus 911 may communicatively couple the components to each other and provide a means for data transfer and flow of control signals between the components. The bus 911 may include one or more of: an internal bus, an address bus, a data bus, a control bus, a controller area network (CAN) bus, an Ethernet bus, a peripheral component interconnect express (PCIe) bus, and any other suitable type of bus that can be implemented in the system 910 to enable communication and coordination between the components within the system 910 in real-time (or near real-time).

[0141] The processor 912 may be implemented in hardware, firmware, or a combination of hardware and software, and may be configured to handle real-time (or near real-time) data processing and control of the control system 910. The processor 912 may include one or more of: a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or another type of processing or computing component that can be implemented in the system 910. In some implementations, the processor 912 may be capable of being programmed to perform one or more operations described herein. Further, the processor 912 may include a plurality of processing units, each of which is dedicatedto performing a specific operation.

[0142] The memory 913 may include one or more mediums for storing temporary data, runtime variables, program instructions, and buffers required for the operations of the control system 910. The memory 913 may include one or more of a flash memory, a read-only memory (ROM), a random-access memory (RAM), a dynamic or static storage device (e.g., a flash memory, a magnetic memory, and / or an optical memory), any other suitable type of memory that can be implemented in the system 910 to store information and / or instructions for use by the processor 912.

[0143] The storage component 914 may be configured to store non-volatile data, such as firmware, configuration settings, calibration data, information, and / or software related to the operation and use of the system 910. For example, the storage component 914 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium, along with a corresponding drive.

[0144] According to example embodiments, the storage component 914 may be configured to store computer-readable or computer-executable instructions for implementing one or more operations of the system 910. The storage component 914 may provide the stored information to the memory 913 for the execution of the processor 912.

[0145] The input component 915 may include one or more input components that permit the system 910 to receive information, such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone). The output component 916may include one or more output components that provide output information from the system 910 (e.g., a display, a speaker, a navigation device, one or more light-emitting diodes (LEDs), etc.) According to example embodiments, the input component 915 and / or the output component 916 may be optional and may be excluded from the system 910.

[0146] The at least one communication interface 917 may include a transceiver-like component (e.g., a transceiver and / or a separate receiver and transmitter) that enables the system 910 to communicate with other components (e.g., ECUs, user devices, etc.), such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. For example, communication interface 917 may include a controller area network (CAN) bus interface, an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, or the like.[00147J According to one or more embodiments, the communication interface 917 may include at least one input / output (I / O) interface, at least one network interface, at least one storage interface, or the like, that enable the components 912-916 to communicate with other components. Further, the communication interface 917 may include one or more application programming interfaces (APIs) that allow the system 910 (or one or more components included therein) to communicate with one or more software applications (e.g., software application deployed in the ECUs, etc.)

[0148] Computer-executable instructions (e.g., software instructions, etc.) may be read into memory 913 and / or storage component 914 from another computer-readable medium or from another device (e.g., a remote server, an external storage, etc.) via, for example, the communicationinterface 917. When executed, the computer-executable instructions stored in memory 913 and / or storage component 914 may cause the processor 912 to perform one or more processes described herein. Additionally, or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more processes described herein. Thus, implementations described herein are not limited to any specific combination of hardware circuitry and software.

[0149] It is contemplated that features, advantages, and significances of example embodiments described hereinabove are merely a portion of the present disclosure, and are not intended to be exhaustive or to limit the scope of the present disclosure. Further descriptions of the features, components, configuration, operations, and implementations of example embodiments of the present disclosure, as well as the associated technical advantages and significances, are provided in the following.[00150J It is understood that the specific order or hierarchy of blocks in the processes / flowcharts disclosed herein is an illustration of example approaches. Based upon design preferences, it is understood that the specific order or hierarchy of blocks in the processes / flowcharts may be rearranged. Further, some blocks may be combined or omitted. The accompanying method claims present elements of the various blocks in a sample order, and are not meant to be limited to the specific order or hierarchy presented.

[0151] Some embodiments may relate to a system, a method, and / or a computer-readable medium at any possible technical detail level of integration. Further, as described hereinabove, one or more of the above components described above may be implemented as instructions stored on a computer readable medium and executable by at least one processor (and / or may include atleast one processor ). The computer-readable medium may include a computer-readable non-transitory storage medium (or media) having computer-readable program instructions thereon for causing a processor (or processors) to carry out operations.

[0152] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0153] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer-readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise coppertransmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0154] Computer readable program code / instructions for carrying out operations may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuitry, or either source code or object code written in any combination of one or more programming languages, including an object-oriented programming languages such as Smalltalk, C++, or the like, and procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a standalone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects or operations.

[0155] These computer readable program instructions may be provided to a processor of a SoC, a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0156] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or another device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0157] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer-readable media according to various embodiments. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). The method,computer system, and computer-readable medium may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in the Figures. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed concurrently or substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.

[0158] It will be apparent that systems and / or methods, described herein, may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting of the implementations. Thus, the operation and behavior of the systems and / or methods were described herein without reference to specific software code-it being understood that software and hardware may be designed to implement the systems and / or methods based on the description herein.

Claims

What is claimed is :

1. A system comprising:a memory storage storing computer-executable instructions; and at least one processor communicatively coupled to the memory storage, wherein the at least one processor is configured to execute the instructions to:obtain a plurality of recorded voice lines of a secondary user;train an artificial intelligence (Al) voice assistant based on the obtained plurality of voice lines of the secondary user; andprovide a generated voice line to a primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the secondary user;wherein the generated voice line is associated with language learning process.

2. The system according to claim 1, wherein the second user comprises a family member of the primary user.

3. The system according to claim 1 , wherein the at least one processor is configured to execute the instructions to obtain the plurality of recorded voice lines of the secondary user by: displaying a plurality of training images associated with a plurality of words to the secondary user;receiving a plurality of voice lines corresponding to the plurality of words associated with the plurality of training images from the secondary user; andstoring the plurality of voice lines received from the secondary user.

4. The system according to claim 3, wherein the plurality of training images comprises at least one predefined training image and at least one training image received from the secondary user.

5. The system according to claim 1 , wherein the at least one processor is configured to execute the instructions to train the Al voice assistant based on the obtained plurality of voice lines of the secondary user by:training a machine learning (ML) model based on the obtained plurality of voice lines of the secondary user; andconfiguring the Al voice assistant based on the trained ML model.

6. The system according to claim 1, wherein the at least one processor is configured to execute the instructions to provide the generated voice line to the primary user using the Al voice assistant by:displaying a learning image associated with a word to the primary user; generating a learning voice line corresponding to the word associated with the learning image using the Al voice assistant; andproviding the generated learning voice line to the primary user, such that the generated learning voice line is provided with the voice of the secondary user.

7. The system according to claim 6, wherein the at least one processor is further configured to execute the instructions to provide the generated voice line to the primary user using the Al voice assistant by:in response to providing the generated learning voice line to the primary user, receiving an input voice line from the primary user;determining whether the input voice line corresponds to the word associated with the learning image;in response to determining that the input voice line corresponds to the word associated with the learning image, generating a positive affirmation voice line using the Al voice assistant, and providing the generated positive affirmation voice line to the primary user, such that the generated positive affirmation voice line is provided with the voice of the secondary user; andin response to determining that the input voice line does not correspond to the word associated with the learning image, providing the generated learning voice line to the primary user, such that the generated learning voice line is provided with the voice of the secondary user.

8. The system according to claim 1, wherein the at least one processor is configured to execute the instructions to provide the generated voice line to the primary user using the Al voice assistant by:displaying a plurality of secondary users who have previously provided the plurality of recorded voice lines;receiving a selection input selecting one of the displayed plurality of secondary users; andproviding the generated voice line to the primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the selected one of the displayed plurality of secondary users.

9. A method comprising:obtaining a plurality of recorded voice lines of a secondary user;training an artificial intelligence (Al) voice assistant based on the obtained plurality of voice lines of the secondary user; andproviding a generated voice line to a primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the secondary user;wherein the generated voice line is associated with language learning process.

10. The method according to claim 9, wherein the second user comprises a family member of the primary user.

11. The method according to claim 9, wherein the obtaining the plurality of recorded voice lines of the secondary user comprises:displaying a plurality of training images associated with a plurality of words to the secondary user;receiving a plurality of voice lines corresponding to the plurality of words associated with the plurality of training images from the secondary user; andstoring the plurality of voice lines received from the secondary user.

12. The method according to claim 11, wherein the plurality of training images comprises at least one predefined training image and at least one training image received from the secondary user.

13. The method according to claim 9, wherein the training the Al voice assistant based on the obtained plurality of voice lines of the secondary user comprises:training a machine learning (ML) model based on the obtained plurality of voice lines of the secondary user; andconfiguring the Al voice assistant based on the trained ML model.

14. The method according to claim 9, wherein the providing the generated voice line to the primary user using the Al voice assistant comprises:displaying a learning image associated with a word to the primary user; generating a learning voice line corresponding to the word associated with the learning image using the Al voice assistant; andproviding the generated learning voice line to the primary user, such that the generated learning voice line is provided with the voice of the secondary user.

15. The method according to claim 14, wherein the providing the generated voice line to the primary user using the Al voice assistant further comprises:in response to providing the generated learning voice line to the primary user, receiving an input voice line from the primary user;determining whether the input voice line corresponds to the word associated with the learning image;in response to determining that the input voice line corresponds to the word associated with the learning image, generating a positive affirmation voice line using the Al voice assistant, and providing the generated positive affirmation voice line to the primary user, such that the generated positive affirmation voice line is provided with the voice of the secondary user; andin response to determining that the input voice line does not correspond to the word associated with the learning image, providing the generated learning voice line to the primary user, such that the generated learning voice line is provided with the voice of the secondary user.

16. The method according to claim 9, wherein the providing the generated voice line to the primary user using the Al voice assistant comprises:displaying a plurality of secondary users who have previously provided the plurality of recorded voice lines;receiving a selection input selecting one of the displayed plurality of secondary users; andproviding the generated voice line to the primary user using the AT voice assistant, such that the generated voice line is provided with a voice of the selected one of the displayed plurality of secondary users.

17. A non-transitory computer-readable recording medium having recorded thereon instructions executable by at least one processor to cause the at least one processor to perform a method comprising:obtaining a plurality of recorded voice lines of a secondary user;training an artificial intelligence (Al) voice assistant based on the obtained plurality of voice lines of the secondary user; andproviding a generated voice line to a primary user using the Al voice assistant, such that the generated voice line is provided with a voice of the secondary user;wherein the generated voice line is associated with language learning process.

18. The non-transitory computer-readable recording medium according to claim 17, wherein the second user comprises a family member of the primary user.

19. The non-transitory computer-readable recording medium according to claim 17, wherein the obtaining the plurality of recorded voice lines of the secondary user comprises:displaying a plurality of training images associated with a plurality of words to the secondary user;receiving a plurality of voice lines corresponding to the plurality of words associated with the plurality of training images from the secondary user; andstoring the plurality of voice lines received from the secondary user.

20. The non-transitory computer-readable recording medium according to claim 19, wherein the plurality of training images comprises at least one predefined training image and at least one training image received from the secondary user.