Intelligent interactive voice device and system with open architecture and voice interaction method
Through an open-architecture intelligent interactive voice device, the story script is dynamically constructed and updated using electronic tags and human-computer interfaces, the problem of insufficient story generation and user interaction in the existing technology is solved, and an interactive experience with high immersion and creative freedom is achieved.
Patent Information
- Application Number
- CN202510194849.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-13
AI Technical Summary
The existing intelligent interactive voice toys have rigidity and lack of immersive participation in story generation and user interaction, and cannot effectively improve users' fun and creative freedom.
An intelligent interactive voice device adopts an open architecture, through the combination of input module, character recognition module, voice playback module and processing module, users can input instructions through electronic tags and human-computer interfaces, dynamically build and update story scripts, and generate corresponding lines and tone characteristics.
It realizes the ability of the user to lead the output of content from the voice device, improves the user's immersive participation experience and interactivity, significantly improves the fun and narrative freedom of the toys, and reduces hardware costs.
Smart Images

Figure CN120148468A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and particularly to an intelligent interactive voice device, system and voice interaction method with an open architecture. Background Art
[0002] With the development and wide popularization of intelligent toys, the interactivity of toys and the demand for users' immersive participation have become the core elements of the product market competitiveness. Therefore, the previously popular story machine voice toys have gradually been replaced by some interactive voice toys.
[0003] Currently, in addition to the conventional storytelling function, this type of intelligent interactive voice toy can also carry out conversations with users, and can also customize stories according to the story scenes and character roles set by users, enabling such stories to break through the versions recorded in books and create rich and imaginative story plots. Therefore, it can arouse users' interest more and improve the playfulness of the toy.
[0004] The disclosure of the above background art content is only used to assist in understanding the concept and technical solution of the present application. It does not necessarily belong to the prior art of the present application, nor will it necessarily give technical guidance; without clear evidence indicating that the above content was publicly available before the filing date of the present application, the above background art should not be used to evaluate the novelty and inventiveness of the present application. Summary of the Invention
[0005] The object of the present invention is to provide an intelligent voice device that further enhances the user's immersive participation experience. An open architecture is adopted to enable users to dominate the content output by the voice device, and the improvement of the interactivity between users and toys greatly enhances the fun of the toy.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] An intelligent interactive voice device with an open architecture includes the following modules:
[0008] An input module configured to receive a first instruction;
[0009] A role recognition module configured to recognize role information corresponding to an external electronic tag;
[0010] A voice playback module configured to play voice information;
[0011] A processing module configured to be electrically connected to the input module, the role recognition module and the voice playback module. The input module sends the received first instruction to the processing module, and the role recognition module sends the recognized role information to the processing module;
[0012] The processing module is configured to construct a story script according to the first instruction, and in response to each time the character recognition module recognizes character information, the processing module generates lines corresponding to the currently recognized character information based on the story script;
[0013] The voice playback module is configured to play voice information corresponding to the currently generated lines.
[0014] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the processing module is further configured to determine a timbre feature adapted to the currently recognized character information and associate it with the currently generated lines;
[0015] The voice information corresponding to the lines played by the voice playback module has a timbre feature associated with the lines.
[0016] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the processing module is further configured to generate narration lines based on the story script and configure a narration timbre feature for the narration lines, where the narration timbre feature is different from the timbre feature corresponding to the lines;
[0017] The voice playback module is further configured to play voice information corresponding to the narration lines with the narration timbre feature.
[0018] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the input module is further configured to receive a second instruction;
[0019] The processing module is configured to update the story script according to the second instruction and generate narration lines and / or lines corresponding to the currently recognized character information based on the updated story script.
[0020] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the story script is updated in the following manner:
[0021] The processing module calls a large language model for context-aware semantic parsing, extracts the modification intention of the second instruction relative to the first instruction, and updates the story script according to the extracted modification intention.
[0022] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, updating the story script includes rewriting the script in at least one of the character layer, background layer, and plot layer, where:
[0023] The rewritten content of the character layer includes at least one of the attributes of the character and the relationship graph between the characters;
[0024] The rewritten content of the background layer includes at least one of time and space migration and environmental detail filling;
[0025] The rewritten content of the plot layer includes at least one of event change and compatibility and extended side plot.
[0026] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the device further includes a communication module, and the processing module is connected to a human - machine interface through the communication module to obtain the first instruction input by the human - machine interface. The first instruction includes one or more pieces of information such as character identity, character personality, character voice, story background, story scene, story style, plot, or scene keywords;
[0027] And / or, the device further includes a voice recognition module configured to recognize the first instruction in voice form. The first instruction includes one or more pieces of information such as character identity, character personality, character voice, story background, story scene, story style, plot, or scene keywords.
[0028] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the device further includes a script - writing database connected to the processing module. The script - writing database is configured to store material data for creating a story script;
[0029] And / or, the processing module is connected to an external storage medium or the cloud through the communication module to obtain material data for creating a story script by accessing the external storage medium or the cloud.
[0030] According to another aspect of the present invention, the present invention provides an open - architecture intelligent interactive voice system, including a plurality of electronic tags and the intelligent interactive voice device as described above.
[0031] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the system further includes a plurality of doll entities corresponding to each electronic tag one by one. Among them, the electronic tag is an NFC tag, a QR code tag, a barcode tag, a Bluetooth beacon, an RFID tag, or a SnapTag.
[0032] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the system further includes a human - machine interaction device with a human - machine interface. The intelligent interactive voice device is connected to the human - machine interaction device to obtain the instruction input by the human - machine interface. The instruction includes one or more pieces of information such as character identity, character personality, character voice, story background, story scene, story style, plot, or scene keywords;
[0033] The human-computer interaction device is integrally provided with the intelligent interactive voice device, or the human-computer interaction device is a smart terminal carrier where a client or application program is connected to the intelligent interactive voice device through a communication module.
[0034] According to another aspect of the present invention, the present invention provides an intelligent voice interaction method with an open architecture, including the following steps:
[0035] Receive a first instruction, where the first instruction includes one or more pieces of information such as character identity, character personality, character voice, story background, story scene, story style, plot, or scenario keywords;
[0036] Construct a story script according to the first instruction;
[0037] Identify the character information corresponding to the external electronic tag;
[0038] Generate lines corresponding to each identified character information based on the story script;
[0039] Play the voice information corresponding to the lines.
[0040] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the intelligent voice interaction method provided by the present invention further includes:
[0041] Generate narration lines based on the story script, and configure a narration tone feature for the narration lines, where the narration tone feature is different from the tone feature of the lines corresponding to the character information; and play the voice information corresponding to the narration lines with the narration tone feature;
[0042] And / or, in response to each identified character information corresponding to the external electronic tag, determine its suitable tone feature, and play the lines corresponding to the current character according to the suitable tone feature.
[0043] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the intelligent voice interaction method provided by the present invention further includes:
[0044] Receive a second instruction, where the second instruction includes one or more pieces of information such as character identity, character personality, character voice, story background, story scene, story style, plot, or scenario keywords;
[0045] Analyze the difference between the second instruction and the first instruction, and extract the modification intention of the second instruction relative to the first instruction;
[0046] Rewrite at least one of the character layer, background layer, and plot layer of the story script according to the modification intention to obtain an updated story script;
[0047] Generate voiceover lines based on the updated story script and / or generate corresponding lines for the recognized character information.
[0048] The beneficial effects brought by the technical solution provided by the present invention are as follows:
[0049] a. Create a story generation system that does not rely on a preset script using an open architecture. The user can dominate the content output by the voice device, enhancing the user's immersive participation experience. The improvement in the interactivity between the user and the toy greatly enhances the fun of the toy.
[0050] b. Realize the intelligent reuse of the character attributes represented by the electronic tag, enabling the user to combine an old toy with an external electronic tag to dynamically activate the AI character and access the dynamic AI narrative ecosystem. While reducing the hardware cost, it significantly improves the narrative freedom and the user's creative participation.
[0051] c. Based on the intelligent voice device, different electronic tags can be combined within a story framework. The combination of electronic tags can break the existing and conventional, and can, according to the user's imagination, construct a brand-new and imaginative story script.
[0052] d. Optimize the multi-character interaction within the script and allow the temporary addition of new characters, enabling intelligent dynamic interpretation of the plot according to the currently recognized electronic tag.
[0053] e. Attach an external electronic tag to an old toy / figurine, reusing the toys purchased in the past and turning the sunk cost into new value. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0055] Figure 1 Schematic block diagram of an intelligent interactive voice device with an open architecture provided for an exemplary embodiment of the present invention;
[0056] Figure 2 Schematic block diagram of an intelligent interactive voice system with an open architecture provided for an exemplary embodiment of the present invention;
[0057] Figure 3 Schematic flowchart of an intelligent voice interaction method with an open architecture provided for an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0059] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0060] The story generation system in typical voice toys often relies on a preset script. Users often cannot trigger the evolution of an open-ended plot and lack the ability to edit the behavior patterns of the characters in the story in real time, making conventional voice toys unable to improve the user experience due to the lack of immersive participation.
[0061] In an embodiment of the present invention, an intelligent interactive voice device with an open architecture is provided, as Figure 1 shown, the intelligent interactive voice device includes the following modules:
[0062] An input module configured to receive a first instruction;
[0063] A role recognition module configured to recognize the role information corresponding to an external electronic tag;
[0064] A voice playback module configured to play voice information;
[0065] A processing module configured to be electrically connected to the input module, the role recognition module, and the voice playback module. The input module sends the received first instruction to the processing module, and the role recognition module sends the recognized role information to the processing module;
[0066] The processing module is configured to construct a story script according to the first instruction, and in response to each time the character recognition module recognizes character information, the processing module generates lines corresponding to the currently recognized character information based on the story script;
[0067] The voice playback module is configured to play voice information corresponding to the currently generated lines.
[0068] See further Figure 1 The processing module can also determine a timbre feature adapted to the currently recognized character information and associate it with the currently generated lines;
[0069] Accordingly, the voice information corresponding to the lines played by the voice playback module has a timbre feature associated with the lines.
[0070] In a specific embodiment, the external electronic tag is an NFC tag, a QR code tag, a barcode tag, a Bluetooth beacon, an RFID tag or a SnapTag. The electronic tag can be physically attached to a physical doll. For example, the tag of Sun Wukong is attached to the Sun Wukong doll, and the tag of Zhu Bajie is attached to the Zhu Bajie doll; alternatively, the electronic tag can be integrated with the doll entity. Users can choose to simply purchase the electronic tag to reuse the old toys / figurines that have been purchased, or they can choose to purchase character products integrated with the electronic tag.
[0071] Attaching the external electronic tag to the old toys / figurines to reuse the toys bought in the past and turning the sunk cost into new value. In this embodiment, there is no need to purchase dedicated interaction devices for each character in the story script. Even the physical doll can adopt a printed two-dimensional plane image or a 3D printed three-dimensional image.
[0072] This embodiment provides an intelligent interactive voice device with an open architecture. The open architecture means a way different from the existing way of realizing interaction according to a fixed mode. This embodiment aims at enabling users to play the role of a director and interact with the intelligent interactive voice device as follows:
[0073] First step: The user issues a first instruction, which can be one or more pieces of information such as character identity, character personality, character timbre, story background, story scene, story style, plot or scenario keywords. There are multiple ways to issue the first instruction: For example, the intelligent interactive voice device is equipped with a microphone that can collect / recognize the user's voice, such as "The four disciples of Tang Seng encountered Ultraman on their journey to the Western Heaven to fetch scriptures and joined hands with him to fight against the monsters". The processing module extracts character identity, story background, story scene, plot or scenario keywords, etc. from the voice information to construct a story script.
[0074] There are also other ways to issue the first instruction. For example, the processing module is communicatively connected to a human-machine interface through the communication module, and the first instruction is input using the human-machine interface. For example, a mobile phone downloads and installs the corresponding App or WeChat mini-program, sets the story background as the journey to the West by Tang Seng on the parameter setting interface, sets the story scene as joining hands with Ultraman to fight against big monsters, sets the plot keywords as the Golden Cudgel, sets the scene keywords as the Flaming Mountains, sets the story style as humorous or second-generation, sets the character identities as Tang Seng, Sun Wukong, Zhu Bajie, Sha Heshang, and Ultraman Zero, and can set / select the personalities and character voice timbres for the above several characters.
[0075] Second step: The user starts to guide the script. Taking the form that the electronic tag is an NFC tag and the electronic tag is physically attached to the doll as an example, correspondingly, the character recognition module of the intelligent interactive voice device is an NFC card reader. The intelligent interactive voice device can be a base. When the doll is brought close to the base, the NFC card reader reads the electronic tag to identify the character information corresponding to the currently selected doll by the user.
[0076] Integrate the large language model and edge computing technology to build an intelligent interactive toy system based on NFC tags. By effectively combining NFC tags with the cloud database, realize the intelligent reuse of the character attributes represented by NFC tags, enabling users to dynamically activate AI characters through the combination of old toys / NFC tags, while reducing hardware costs and significantly improving the narrative freedom and user creation participation.
[0077] In this embodiment, wait for and identify the information interaction of the NFC tag, and send the unique serial number of the NFC tag to the cloud through, for example, the WIFI network. Obtain data such as the story plot background and character information related to the character through this serial number, providing a basis for subsequent story interpretation.
[0078] Third step: The processing module generates lines corresponding to the current character information based on the constructed story script. For example, when the user first selects the NFC tag corresponding to Sun Wukong and brings it close to the NFC card reader, the processing module generates the lines of Sun Wukong.
[0079] Based on the background character information and user input in the current story script, use the large language model specialized in the story field, and combine the story content to generate the dialogue text data of the character corresponding to the NFC tag at the current stage.
[0080] Fourth step: The voice playback module plays the voice information corresponding to the currently generated lines: According to the timbre information of the character corresponding to the NFC tag, using the voice synthesis technology, the system generates the character voice of the specified timbre, interacts with the user, and promotes the development of the story plot.
[0081] In a specific embodiment, the voice timbre of Sun Wukong can be set by a first instruction, such as selecting a recommended classic Wukong timbre; in other embodiments, the voice timbre corresponding to the character can be determined by reading an electronic tag, and then the voice playback module outputs lines with corresponding timbre characteristics, such as "Hey, where did you come from? You look powerful! But I, Sun Wukong, have never been afraid of anyone on this journey to the West! Are you here to help or to make trouble? If you are here to help, that's great, as this monster is worried that no one will help it deal with it; if you are here to make trouble, don't blame my golden hoop for being blind! But looking at your figure, you seem like a serious person. First, tell me who you are and why you are here!"
[0082] Since the character attributes can be set / selected for the character in the first step, users can break through the conventional character image, such as changing Sun Wukong to be humble, and changing Zhu Bajie's character to be hardworking and enthusiastic, and can change the story script and the corresponding character lines. This enables the intelligent interactive voice device to respond to the user's wild imagination and output a new and imaginative story script that matches it, realizing the one card of the electronic tag with many faces.
[0083] Then the user repeats the third and fourth steps. This time, the user selects the Ultraman NFC tag and brings it close to the NFC card reader. Then the voice playback module plays the lines with Ultraman's voice characteristics: "Haha, you monkey are such an interesting character! I am Ultraman, from a distant universe. My mission is to protect the peace of the universe and eliminate all evil forces. This time I sensed that there are dark forces at work on this land, so I came here specially. I didn't expect to meet you here. It seems that we are like-minded partners!"
[0084] Multi-character interaction does not rely on preset scripts, so there is no limit on the order in which users place the NFC tags corresponding to each character. On the contrary, the story script will be dynamically adjusted according to the order in which users place the NFC tags, making the character data and script information storage dynamic, and dynamically advancing the development of the story based on user interaction.
[0085] In a specific embodiment, the processing module is further configured to generate narration lines based on the story script and configure a narration voice characteristic for the narration lines, where the narration voice characteristic is different from the voice characteristic corresponding to the lines; the voice playback module plays the narration lines with the narration voice characteristic: "Ultraman said, turning his gaze to the monster in the distance, and his tone became serious", making the transition of the performances of each character in the entire script natural; then the voice playback module then plays the unfinished lines of Ultraman, and at this time, it can be in a slightly serious style: "That monster comes from the dark universe. It has powerful strength. It may be very difficult for you alone to deal with it. Sun Wukong, let's cooperate! For justice and peace, fight together!"
[0086] Then the user brings the NFC tag of Tang Seng close to the NFC reader again. The processing module generates a connecting narration line related to Tang Seng and the lines of Tang Seng. Then the voice playback module first plays with the narration voice characteristic: "Tang Seng saw the conversation between Sun Wukong and Ultraman and realized that both sides were here to help subdue demons and eliminate monsters. So he quickly stepped forward, put his hands together, smiled, and said:", and the addition of the narration makes the story script heard by the user more complete and natural; then it switches to the voice characteristic of Tang Seng's voice to play: "Amitabha, goodness! Benefactor Ultraman, Sun Wukong has an impatient temper and may have offended you with his words. Please forgive me. The four of us master and disciples are on a journey to the Western Heaven to obtain the scriptures. Along the way, there have been many monsters obstructing us. Today we have come to the Flaming Mountains. Fortunately, we have the help of the benefactor. This is truly a blessing for us. Since the benefactor has come to protect the peace of the universe, then we are people of the same path. Since this is the case, please discuss with my disciples how to jointly subdue that monster to ensure the peace of this place." In a specific embodiment, in addition to the voice characteristic, each character can also have corresponding voice characteristics and speech rate characteristics in different emotions, making the voice broadcast obtained by the user more vivid and vivid.
[0087] During the process of the plot performance, the user instructions will continue to be monitored. For example, as the need for the plot to progress, the user can issue a second instruction again as in the first step, which can be one or more pieces of information such as character identity, character personality, character voice, story background, story scene, story style, plot or scene keywords. The processing module is configured to update the story script according to the second instruction and generate narration lines and / or lines corresponding to the currently recognized character information based on the updated story script.
[0088] For example, after the voice playback of the character of Tang Seng is completed, the user says: "Just as Tang Seng's voice fell, a monster walked into view with heavy steps". The processing module calls the large language model to perform context-aware semantic parsing, extracts the modification intention of the second instruction relative to the first instruction, and updates the story script according to the extracted modification intention, such as rewriting the script in at least one of the character layer, background layer, and plot layer of the script, where:
[0089] The rewritten content of the character layer includes at least one of the attributes of the characters and the relationship graph among the characters. In this embodiment, the monster will be added to the relationship graph of the characters;
[0090] The rewritten content of the background layer includes at least one of space-time migration and environmental detail filling. In this embodiment, the spatial shot will change from far to near, and the architectural details of the road surface walked by the monster can also be filled;
[0091] The rewritten content of the plot layer includes at least one of event change and compatibility, and extended side plot. In this embodiment, due to the appearance of the monster, the plot will turn to the event that Tang Seng and his disciples fight against the monster in cooperation with Ultraman.
[0092] During the process of updating the story script based on the modification intention, priority determination is required, that is, conflicting instructions are processed, and associated knowledge can be obtained from some existing semantic retrieval systems such as Wikidata, ConceptNet, etc., and logical verification is performed to ensure the coherence of the rewritten script. For example, the role of a modern doctor cannot be directly transplanted into an ancient background and needs to be automatically adjusted to a "herbalist".
[0093] For example, when the user brings the NFC tag of the monster close to the NFC reader, the voice playback module can first play the environmental background sounds (such as earthquake, crowd screaming and other sound effects) imitating the actions of the monster, and play "I am the master of this land, the incarnation of darkness - the Dark Troll! Come on, stupid guys! If you kneel down and beg for mercy now, I can consider making your death more painless. Otherwise, I will make your life worse than death!" in the voice timbre characteristics of the monster.
[0094] Subsequently, according to the order of the characters corresponding to the NFC tags brought close by the user, the intelligent interactive voice device outputs the narration, sound effects of the battle scene and the lines of the corresponding characters. Different orders of the characters will affect different script interpretations, and the user can add several second instructions at any time during the script interpretation to gradually update the story script step by step. The intelligent interactive voice device can determine the direction of the story script until the end according to the second instructions of the user.
[0095] The above intelligent interactive voice device is actually an artificial intelligence large model with a script creation database. Based on the Natural Language Understanding (NLU) module, it uses a large language model to perform context-aware semantic parsing. The script creation database is connected to the processing module, and the script creation database is configured to store material data for creating story scripts. In another embodiment, the processing module is connected to an external storage medium or the cloud through a communication module to obtain material data for creating story scripts by accessing the external storage medium or the cloud. The present invention does not limit the form in which the material data for creating story scripts is obtained.
[0096] See Figure 2 , an embodiment of the present invention provides an intelligent interactive voice system with an open architecture, including a plurality of electronic tags and the intelligent interactive voice device as described above; the system further includes a human-computer interaction device with a human-machine interface, and the intelligent interactive voice device is connected to the human-computer interaction device to obtain instructions input through the human-machine interface. The instructions include one or more pieces of information such as character identity, character personality, character tone color, story background, story scene, story style, and plot keywords;
[0097] The human-computer interaction device is integrally provided with the intelligent interactive voice device, or the human-computer interaction device is an intelligent terminal carrier where a client or application program is connected to the intelligent interactive voice device through a communication module.
[0098] See Figure 3 , an embodiment of the present invention provides an intelligent voice interaction method with an open architecture, including the following steps:
[0099] Receive a first instruction, where the first instruction includes one or more pieces of information such as character identity, character personality, character tone color, story background, story scene, story style, and plot keywords; specifically, collect the user's voice through a microphone and extract relevant information therefrom to generate the first instruction;
[0100] Construct a story script according to the first instruction;
[0101] Identify the character information corresponding to the external electronic tag;
[0102] Generate lines corresponding to each identified character information based on the story script;
[0103] Play the voice information corresponding to the lines.
[0104] The intelligent voice interaction method provided by the embodiment of the present invention further includes: generating a narrator's lines based on the story script, and configuring a narrator's voice characteristic for the narrator's lines, where the narrator's voice characteristic is different from the voice characteristic of the lines corresponding to the character information; and playing the voice information corresponding to the narrator's lines with the narrator's voice characteristic.
[0105] And / or, in response to each recognition of the character information corresponding to the external electronic tag, determining its adapted voice characteristic, and playing the lines corresponding to the current character according to the adapted voice characteristic.
[0106] The intelligent voice interaction method provided by the embodiment of the present invention further includes: receiving a second instruction, where the second instruction includes one or more pieces of information such as character identity, character personality, character voice, story background, story scene, story style, and plot keywords.
[0107] Analyzing the difference between the second instruction and the first instruction, extracting the modification intention of the second instruction relative to the first instruction; according to the modification intention, rewriting at least one of the character layer, background layer, and plot layer of the story script to obtain an updated story script; generating narrator's lines based on the updated story script, and / or generating corresponding lines for the recognized character information.
[0108] The intelligent voice interaction method provided by this embodiment and the intelligent interactive voice device provided by the above embodiment belong to the same inventive concept. Here, the entire content of the embodiment of the intelligent interactive voice device is incorporated into the embodiment of this intelligent voice interaction method by reference, and will not be elaborated herein.
[0109] The present invention constructs an open architecture, dynamic, and personalized voice interaction system through an innovative multi-agent collaboration architecture and NFC interaction technology, providing users with a creative and interactive story experience. The present invention comprehensively elaborates the technical architecture, key implementation mechanisms, and innovative values of the system, not only breaking through the limitations of traditional interaction modes, but also providing users with a great deal of creative freedom, effectively solving the core problems such as rigid narration, high hardware costs, and single interaction in the prior art, providing users with a low-cost, high-degree-of-freedom, and strong immersion interactive experience, and having significant technological progress and market application value.
[0110] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0111] The above are only specific embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. An open-architecture intelligent interactive voice device, characterized in that: Includes the following modules: An input module configured to receive a first instruction; A role identification module, which is configured to identify role information corresponding to the external electronic tag; A voice playing module, which is configured to play voice information; a processing module, which is configured to be electrically connected to the input module, the role recognition module and the voice playback module, wherein the input module sends the received first instruction to the processing module, and the role recognition module sends the recognized role information to the processing module; The processing module is configured to construct a story script according to the first instruction, and in response to each time the character identification module identifies the character information, the processing module generates a line corresponding to the currently identified character information based on the story script; The voice playing module is configured to play the voice information corresponding to the currently generated lines.
2. The intelligent interactive voice device according to claim 1, characterized in that: The processing module is further configured to determine a timbre feature that matches the currently recognized character information and associate it with the currently generated lines; The speech information corresponding to the lines played by the speech playing module has timbre characteristics associated with the lines.
3. The intelligent interactive voice device according to claim 1, characterized in that: The processing module is further configured to generate narration lines based on the story script, and configure narration timbre features for the narration lines, wherein the narration timbre features are different from timbre features corresponding to the lines; The voice playing module is also configured to play voice information corresponding to the narration lines having narration timbre characteristics.
4. The intelligent interactive voice device according to claim 1, characterized in that: The input module is further configured to receive a second instruction; The processing module is configured to update the story script according to the second instruction, and generate narration lines and / or lines corresponding to the currently recognized character information based on the updated story script.
5. The intelligent interactive voice device according to claim 4, characterized in that: Update the story script by: The processing module calls the large language model to perform context-aware semantic analysis, extracts the modification intention of the second instruction relative to the first instruction, and updates the story script according to the extracted modification intention.
6. The intelligent interactive voice device according to claim 4, characterized in that: The updating of the story script includes rewriting the script at least one of the character layer, the background layer, and the plot layer, wherein: The rewritten content of the role layer includes at least one of the attributes of the role and the relationship map between the roles; The rewritten content of the background layer includes at least one of time-space migration and environmental detail filling; The rewritten content of the plot layer includes at least one of event changes and compatibility, and expansion of branch plots.
7. The intelligent interactive voice device according to claim 1, characterized in that: The device further comprises a communication module, and the processing module is connected to the human-machine interface through the communication module to obtain the first instruction input by the human-machine interface, wherein the first instruction comprises one or more information of character identity, character personality, character timbre, story background, story scene, story style, plot or scenario keywords; And / or, the device also includes a speech recognition module, which is configured to recognize a first instruction in speech form, wherein the first instruction includes one or more information of character identity, character personality, character timbre, story background, story scene, story style, plot or situation keywords.
8. The intelligent interactive voice device according to claim 1, characterized in that: The device also includes a script creation database, which is connected to the processing module, and the script creation database is configured to store material data for creating story scripts; And / or, the processing module is connected to an external storage medium or a cloud via a communication module, so as to obtain material data for creating a story script by accessing the external storage medium or the cloud.
9. An open-architecture intelligent interactive voice system, characterized in that: The invention comprises a plurality of electronic tags and an intelligent interactive voice device as claimed in any one of claims 1 to 8.
10. The intelligent interactive voice system according to claim 9, characterized in that: The system also includes a plurality of doll entities, which correspond one to one with respective electronic tags, wherein the electronic tags are NFC tags, QR code tags, barcode tags, Bluetooth beacons, RFID tags or SnapTags.
11. The intelligent interactive voice system according to claim 9, characterized in that: The system further comprises a human-machine interaction device having a human-machine interface, the intelligent interactive voice device being connected to the human-machine interaction device to obtain instructions inputted by the human-machine interface, the instructions comprising one or more information of character identity, character personality, character timbre, story background, story scene, story style, plot or scenario keywords; The human-computer interaction device is integrated with the intelligent interactive voice device, or the human-computer interaction device is a smart terminal carrier where a client or application program is located and is connected to the intelligent interactive voice device via a communication module.
12. An open-architecture intelligent voice interaction method, characterized in that: The following steps are involved: receiving a first instruction, wherein the first instruction includes one or more information of a character identity, a character personality, a character voice, a story background, a story scene, a story style, a plot, or a scenario keyword; Constructing a story script according to the first instruction; Identify the role information corresponding to the external electronic tag; Generate lines corresponding to each identified character information based on the story script; Play the voice information corresponding to the lines.
13. The intelligent voice interaction method according to claim 12, characterized in that: Also includes: Generate narration lines based on the story script, and configure narration timbre features for the narration lines, wherein the narration timbre features are different from timbre features of lines corresponding to the character information; and playing the voice information corresponding to the narration lines with the narration timbre characteristics; And / or, in response to each recognition of the role information corresponding to the external electronic tag, the corresponding timbre characteristics are determined, and the lines corresponding to the current role are voice-played according to the corresponding timbre characteristics.
14. The intelligent voice interaction method according to claim 12, characterized in that: Also includes: receiving a second instruction, the second instruction including one or more information of character identity, character personality, character timbre, story background, story scene, story style, plot or scenario keywords; Analyze the difference between the second instruction and the first instruction, and extract the modification intention of the second instruction relative to the first instruction; According to the modification intention, at least one of the character layer, the background layer, and the plot layer of the story script is rewritten to obtain an updated story script; Generate narration lines based on the updated story script, and / or generate corresponding lines for the identified character information.
Citation Information
Cited By
Large model AI interaction system based on master-slave system
CN120723198A