Human-computer interaction method and device, equipment, storage medium and program product

By receiving voice information and utilizing vertical domain and topic recognition models, as well as voiceprint and acoustic feature recognition technologies, the system automatically determines whether the user is a child and outputs a child-styled response. This solves the problem of insufficient intelligence in existing technologies and improves the intelligence and user experience of human-computer interaction for children.

CN121918786APending Publication Date: 2026-04-24SHANGHAI LIXIANG AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI LIXIANG AUTOMOBILE CO LTD
Filing Date
2024-10-24
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies cannot proactively determine whether a child-styled response is needed in human-computer interaction scenarios involving children, resulting in low intelligence.

Method used

By receiving voice information input by the user, the system determines the vertical domain or topic to which the voice content belongs, and judges whether the user is a child based on the voice information. If the user is a child, the system outputs a child-style response. The system uses technologies such as vertical domain recognition model, topic recognition model, voiceprint and acoustic feature recognition to achieve automatic judgment.

Benefits of technology

It enables automatic determination of whether a child-styled response is needed in human-computer interaction scenarios, improving the intelligence of human-computer interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121918786A_ABST
    Figure CN121918786A_ABST
Patent Text Reader

Abstract

The invention discloses a man-machine interaction method and device, equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence, and the method comprises the steps: after receiving voice information inputted by a user, determining a vertical domain or topic to which the voice content of the voice information belongs, and if the vertical domain to which the voice content belongs is a target vertical domain needing children stylized reply or the topic is a target vertical domain needing children stylized reply; the topic to which the voice content belongs is a target topic needing children stylized reply, and judging whether the user is a children user based on the voice information; and if the user is a child user, outputting child stylized reply information generated based on the voice content. After the voice information input by the user is received, when it is determined that the child stylized reply condition is met according to the voice information and the voice content of the voice information, the user is replied in the child style, whether child stylized reply is needed or not is automatically judged, and the intelligence of man-machine interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a human-computer interaction method, apparatus, device, storage medium, and program product. Background Technology

[0002] With the development of artificial intelligence technology, human-computer interaction is increasingly being applied in various scenarios, such as knowledge Q&A, intelligent vehicle systems, and virtual human interactive screens.

[0003] Currently, in human-computer voice interaction scenarios with children, the machine will only provide a child-style response when the user requests it. It cannot proactively determine whether a child-style response is needed, resulting in low intelligence. Summary of the Invention

[0004] In view of the above problems, this application provides a human-computer interaction method, apparatus, device, storage medium, and program product to improve the intelligence of human-computer interaction. The specific solution is as follows:

[0005] The first aspect of this application provides a human-computer interaction method, including:

[0006] Receive voice information input by the user;

[0007] Determine the vertical domain or topic to which the voice content of the voice information belongs;

[0008] If the vertical domain to which the voice content belongs is the target vertical domain that requires a child-styled response, or if the topic to which the voice content belongs is the target topic that requires a child-styled response, determine whether the user is a child user based on the voice information;

[0009] If the user is a child, output a child-styled response message generated based on the voice content.

[0010] In one possible implementation, determining the vertical domain or topic to which the speech content of the speech information belongs includes:

[0011] The speech content is processed by a vertical domain recognition model to obtain the vertical domain to which the speech content belongs;

[0012] Alternatively, the speech content can be processed using a topic recognition model to obtain the topic to which the speech content belongs.

[0013] In one possible implementation, determining whether the user is a child user based on the voice information includes:

[0014] Extract the voiceprint of the speech information;

[0015] If a registered voiceprint that matches the voiceprint exists, obtain the age information corresponding to the registered voiceprint;

[0016] If the age information is within a preset range, the user is determined to be a child user;

[0017] If no registered voiceprint matches the voiceprint, extract the acoustic features of the speech information;

[0018] The acoustic features are input into the growth stage recognition model to obtain the recognition result output by the growth stage recognition model, and the recognition result indicates whether the user is a child.

[0019] In one possible implementation, determining whether the user is a child user based on the voice information includes:

[0020] Extract the acoustic features of the speech information;

[0021] The acoustic features are input into the growth stage recognition model to obtain the recognition result output by the growth stage recognition model, and the recognition result indicates whether the user is a child.

[0022] In one possible implementation, the output, based on child-styled response information generated from the speech content, includes:

[0023] Based on the voice content, one or more child-styled response segments are generated; wherein, when multiple child-styled response segments are generated, adjacent child-styled response segments are connected by transition information, wherein the transition information is a summary of the latter child-styled response segment in the two adjacent child-styled response segments.

[0024] Output one or more paragraphs of child-styled response content.

[0025] In one possible implementation, each child-styled response satisfies at least one of the following:

[0026] The wording should be childlike;

[0027] The length is less than the preset number of characters;

[0028] The sentence structure uses a simple sentence.

[0029] In one possible implementation, the output, based on child-styled response information generated from the speech content, includes:

[0030] Determine whether multiple rounds of interaction are needed for the voice content;

[0031] Based on the judgment result, output a child-styled response message generated from the spoken content; specifically:

[0032] If the determination result is yes, a child-styled first response content and a first question are generated based on the voice content. The first response content is a part of the target response content corresponding to the voice content, and the answer to the first question is related to the remaining content of the target response content. After outputting the first response content and the first question, the process returns to the step of receiving voice information input by the user and its subsequent steps to generate and output the remaining content. Alternatively, a child-styled target response content and a second question are generated based on the voice content. The second question is used to guide the child user to think about the related content of the voice content, and the related content is different from the target response content. After outputting the target response content and the second question, the process returns to the step of receiving voice information input by the user and its subsequent steps to generate and output the related content.

[0033] If the judgment result is negative, generate the target response content in a child-style based on the voice content, and output the target response content.

[0034] In one possible implementation, the first response content and the first question generated based on the voice content, the remaining content, the target response content and the second question, and the related content satisfy at least one of the following:

[0035] The wording used is childlike;

[0036] The length is less than the preset number of characters;

[0037] The sentence structure uses a simple sentence.

[0038] In one possible implementation, the determination of whether multi-turn interaction is required for the voice content includes:

[0039] The speech content is subjected to semantic understanding to obtain a semantic understanding result; the semantic understanding result indicates whether multi-turn interaction is required for the speech content.

[0040] Alternatively, semantic understanding can be performed on the speech content and historical interaction content to obtain a semantic understanding result; this semantic understanding result indicates whether multiple rounds of interaction are required for the speech content.

[0041] Alternatively, semantic understanding can be performed on the speech content, or semantic understanding can be performed on the speech content and historical interaction content; if the semantic understanding result indicates that the user accepts multi-turn interaction for the speech content, the speech content is processed to obtain a processing result; the processing result indicates whether multi-turn interaction for the speech content is required.

[0042] In one possible implementation, the output, based on child-styled response information generated from the speech content, includes:

[0043] Output a child-styled reply message generated based on the voice content and the user's nickname; the reply message carries the user's nickname.

[0044] A second aspect of this application provides a human-computer interaction device, comprising:

[0045] The receiving module is used to receive voice information input by the user;

[0046] The determination module is used to determine the vertical domain or topic to which the voice content of the voice information belongs;

[0047] The judgment module is used to determine whether the user is a child user based on the voice information if the vertical domain to which the voice content belongs is a target vertical domain that requires a child-styled response, or if the topic to which the voice content belongs is a target topic that requires a child-styled response.

[0048] The output module is used to output child-styled response information generated based on the voice content if the user is a child user.

[0049] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the human-computer interaction method described in the first aspect or any implementation thereof.

[0050] A fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:

[0051] The memory is used to store computer programs;

[0052] The processor is used to execute the computer program so that the electronic device can implement the human-computer interaction method of the first aspect or any implementation thereof.

[0053] The fifth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to perform a human-computer interaction method as described in the first aspect or any implementation thereof.

[0054] By utilizing the above technical solutions, this application discloses a human-computer interaction method, apparatus, device, storage medium, and program product. After receiving voice information input by a user, the method determines the vertical domain or topic to which the voice content belongs. If the vertical domain to which the voice content belongs is a target vertical domain requiring a child-style response, or if the topic to which the voice content belongs is a target topic requiring a child-style response, the method determines whether the user is a child based on the voice information. If the user is a child, the method outputs child-style response information generated based on the voice content. This application, after receiving voice information input by a user, determines whether the conditions for a child-style response are met based on the voice information and its content (i.e., the user-input voice content belongs to a target vertical domain or target topic requiring a child-style response, and the user is a child), and responds to the user in a child-style manner. This achieves automatic determination of whether a child-style response is needed, improving the intelligence of human-computer interaction. Attached Figure Description

[0055] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0056] Figure 1 A flowchart illustrating an implementation of the human-computer interaction method provided in this application;

[0057] Figure 2 A flowchart illustrating an implementation of this application for determining whether a user is a child based on voice information;

[0058] Figure 3 This application provides an alternative implementation flowchart for determining whether a user is a child based on voice information.

[0059] Figure 4 A flowchart illustrating an implementation of outputting stylized child-based response information generated from speech content, as provided in this application;

[0060] Figure 5 A flowchart illustrating an implementation of this application for generating child-styled response information based on voice content according to the judgment result;

[0061] Figure 6 A schematic diagram of a human-computer interaction device provided in this application;

[0062] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0063] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0064] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0065] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0066] The human-computer interaction method provided in this application can be used in terminal devices, which may include, but are not limited to, any of the following: mobile phones, tablets, laptops, computers, all-in-one machines, in-vehicle devices, etc.

[0067] like Figure 1 The diagram shown is a flowchart of one implementation of the human-computer interaction method provided in this application, which may include:

[0068] Step S101: Receive voice information input by the user.

[0069] The voice information input by the user can be entered according to the user's own needs, or it can be entered in response to a question posed by the human-computer interaction device.

[0070] Users here can be either children or adults.

[0071] Step S102: Determine the vertical domain or topic to which the voice content of the voice information belongs.

[0072] It can perform speech recognition on voice information to obtain the voice content; then perform vertical domain recognition or topic recognition on the voice content to obtain the vertical domain or topic to which the voice content belongs.

[0073] Vertical domain refers to a specific field; topics are a further subdivision of a vertical domain, meaning that a vertical domain can include multiple topics.

[0074] Step S103: If the vertical domain to which the voice content of the voice information belongs is the target vertical domain that requires a child-styled response, or if the topic to which the voice content of the voice information belongs is the target topic that requires a child-styled response, determine whether the user is a child user based on the voice information.

[0075] Optionally, which verticals or topics require child-style responses are pre-defined. Based on this, a vertical whitelist or topic whitelist can be set. If the vertical to which the voice content belongs is in the vertical whitelist, it means that the vertical to which the voice content belongs is the target vertical to which a child-style response is required; otherwise, it means that the vertical to which the voice content belongs is the vertical to which a child-style response is not required. Similarly, if the topic to which the voice content belongs is in the topic whitelist, it means that the topic to which the voice content belongs is the target topic to which a child-style response is required; otherwise, it means that the topic to which the voice content belongs is the topic to which a child-style response is not required.

[0076] Optionally, the following method can be used to determine whether a user is a child user: extract the voiceprint of the voice information; if there is a registered voiceprint that matches the extracted voiceprint, obtain the age information corresponding to the registered voiceprint; if the obtained age information is within a preset range, determine that the user is a child user; otherwise, determine that the user is not a child user.

[0077] Optionally, the system can collect and process the user's image to identify whether the user is a child.

[0078] Step S104: If the user is a child, output a child-styled response message generated based on the voice content.

[0079] If the user is a child, a child-style response is required. Therefore, a child-style response message is generated and output.

[0080] In this embodiment of the application, the child-styled response information can be generated and output at once, or it can be generated and output multiple times in a multi-round interaction manner.

[0081] The human-computer interaction method provided in this application, after receiving voice information input by the user, determines whether the conditions for a child-styled response are met based on the voice information and its content. That is, when the voice content input by the user belongs to the target vertical domain or target topic that requires a child-styled response, and the user is a child, the method replies to the user in a child-style manner. This achieves automatic judgment on whether a child-styled response is needed, thereby improving the intelligence of human-computer interaction.

[0082] In an optional embodiment, one way to implement the above-mentioned determination of the vertical domain to which the voice content of the voice information belongs can be:

[0083] The speech content is processed by a vertical domain recognition model to obtain the vertical domain to which the speech content belongs.

[0084] A vertical domain recognition model can be trained using text as training samples and the vertical domain to which the text belongs as sample labels. The specific training process can be as follows: input the training samples into the vertical domain recognition model to obtain the vertical domain recognition results output by the model. With the goal of making the vertical domain recognition results as close as possible to the labels of the training samples, the parameters of the vertical domain recognition model are updated. That is, compared with the vertical domain recognition results obtained by processing the training samples using the unupdated vertical domain recognition model, the vertical domain recognition results obtained by processing the training samples using the updated vertical domain recognition model are closer to the labels of the training samples.

[0085] By using a vertical domain recognition model to determine the vertical domain to which the speech content belongs, automatic recognition of the vertical domain to which the speech content belongs is achieved.

[0086] In an optional embodiment, one way to implement the above-mentioned determination of the topic to which the voice content of the voice information belongs can be:

[0087] The topic recognition model is used to process the speech content and determine the topic to which the speech content belongs.

[0088] A topic recognition model can be trained using text as training samples and the topics to which the text belongs as sample labels. The specific training process can be as follows: input the training samples into the topic recognition model to obtain the topic recognition results output by the model; update the parameters of the topic recognition model with the goal of making the topic recognition results as close as possible to the labels of the training samples. That is, compared with the topic recognition results obtained by processing the training samples using the unupdated topic recognition model, the topic recognition results obtained by processing the training samples using the updated topic recognition model are closer to the labels of the training samples.

[0089] By using a topic recognition model to determine the topic to which the audio content belongs, automatic topic recognition of the audio content is achieved.

[0090] In an optional embodiment, the flowchart of one method for determining whether a user is a child based on voice information is as follows: Figure 2 As shown, it may include:

[0091] Step S201: Extract the voiceprint of the speech information.

[0092] Step S202: If a registered voiceprint that matches the voice information exists, obtain the age information corresponding to the registered voiceprint.

[0093] In this application, users can register voiceprint information on their terminal devices and fill in personal information corresponding to the registered voiceprint information, such as nickname, and may also include, but is not limited to, at least one of the following information: gender, age, job title, etc.

[0094] Optionally, for any of the above information, if the user does not fill it in, the field can be preset information. For example, if the user does not fill in their age, the age will be 0.

[0095] The voiceprint of the voice information can be compared with each of the registered voiceprints to determine whether there is a registered voiceprint that matches the voice information.

[0096] Step S203: If the obtained age information is within the preset range, determine that the user is a child user.

[0097] As an example, the preset range can be greater than 0 and less than or equal to 8 years old.

[0098] If the obtained age information is not within the preset range, it is determined that the user is not a child user.

[0099] Step S204: If there is no registered voiceprint that matches the voiceprint of the speech information, extract the acoustic features of the speech information.

[0100] If a user has not registered their voiceprint information, their age information cannot be obtained. In this case, it is possible to determine whether the user is a child based on the acoustic characteristics of their voice information.

[0101] The acoustic features can be at least one of the following: fundamental frequency features (Pitch), Mel-scale frequency cepstral coefficients (MFCC), or Mel-scale filter bank (Fbank).

[0102] Step S205: Input the acoustic features into the growth stage recognition model to obtain the recognition result output by the growth stage recognition model. This recognition result indicates whether the user is a child.

[0103] A developmental stage recognition model can be trained using speech as training samples and the developmental stage (child or adult) of the speaker as the sample label. The specific training process can be as follows: input the training samples into the developmental stage recognition model to obtain the developmental stage recognition result output by the model; update the parameters of the developmental stage recognition model with the goal of making the developmental stage recognition result as close as possible to the label of the training samples. That is, compared with the developmental stage recognition result obtained by processing the training samples using the unupdated model, the developmental stage recognition result obtained by processing the training samples using the updated model is closer to the label of the training samples.

[0104] In the above embodiments, the user's developmental stage is identified by combining voiceprint and acoustic features. In another embodiment, the user's developmental stage can be identified solely based on the acoustic features of the voice. Based on this, the flowchart for another implementation of determining whether a user is a child based on voice information is as follows: Figure 3 As shown, it may include:

[0105] Step S301: Extract the acoustic features of the speech information.

[0106] The acoustic features can be at least one of the following: fundamental frequency features (Pitch), Mel-scale frequency cepstral coefficients (MFCC), or Mel-scale filter bank (Fbank).

[0107] Step S302: Input the acoustic features into the growth stage recognition model to obtain the recognition result output by the growth stage recognition model. This recognition result indicates whether the user is a child.

[0108] For details on the specific implementation, please refer to the aforementioned embodiments, which will not be repeated here.

[0109] This application enables automatic identification of whether a user is a child user based on the acoustic characteristics of the user's registered voiceprint or the user's voice.

[0110] This application research found that for child users, complex questions often require lengthy responses, and directly outputting the response content can cause them to lose focus. To maintain child users' attention, this application proposes dividing the response content into multiple segments. After outputting one segment, and before outputting the next segment, a summary of the next segment is output, giving the child user a general understanding and thus helping them maintain attention. Based on this, in an optional embodiment, one implementation of the above-mentioned output of child-styled response information generated from speech content can be:

[0111] Generate one or more child-styled response messages based on the voice content.

[0112] If the reply message only includes a short, child-style reply, it means that the user's question is relatively simple (i.e., the answer content is relatively short). This short, child-style reply is all the reply content corresponding to the voice content, and can be output directly.

[0113] If the response information includes at least two child-style response pieces, and there is also transition information between the two adjacent child-style response pieces, the transition information is a summary of the latter child-style response piece in the two adjacent child-style response pieces.

[0114] If the response includes at least two child-style responses, it indicates that the user's question is relatively complex (the answer is lengthy), and the response needs to be output in multiple segments. When outputting the response in segments, after each child-style response segment, output a brief transition message. This transition message uses concise language to tell the user what the next content will be, so that the child user can look forward to the following content and thus help maintain the child user's attention.

[0115] Output one or more child-styled response segments. When outputting these child-styled response segments, output a transition message after each segment before outputting the next segment.

[0116] Furthermore, each child-styled response must meet at least one of the following criteria:

[0117] The wording should be childlike;

[0118] The length is less than the preset word count;

[0119] The sentence structure uses a simple sentence.

[0120] When each child-styled response meets at least one of the above criteria, it makes it easier for child users to understand and accept the response information, resulting in a smoother human-computer interaction.

[0121] In an optional embodiment, one implementation of generating one or more child-styled response segments based on voice content can be:

[0122] Adding the target information to the instruction template yields the instruction (prompt). The target information includes the voice content of the speech information and historical interaction information. The instruction prompts the large model to respond to the questions posed by children in a child-friendly style.

[0123] In the first round of interaction, the historical interaction information is empty. In the case of at least two rounds of interaction, the historical interaction information refers to the user input information and the response information output by the human-computer interaction device before the current round.

[0124] Input the instructions into the large model to get one or more child-styled responses output by the large model.

[0125] In this embodiment, it is not necessary to determine whether the user's question (i.e., the voice content) is a complex question or a simple question. By fine-tuning the large model through training data, the large model can understand whether the user's question is a simple question or a complex question, and then output at least one corresponding child-style response.

[0126] The training dataset for fine-tuning the large model in this application consists of several interaction data points. Each interaction data point represents one round of interaction data, which includes: a user question and a response to the user question. If the user question is simple, the response is a single paragraph, representing the complete answer to the user question. If the user question is complex, the response consists of multiple paragraphs, plus transition information between adjacent paragraphs. Each paragraph is part of the complete answer to the user question, and the transition information between adjacent paragraphs summarizes the latter paragraph, informing the child user what the next output should be. The transition information between adjacent paragraphs can be merged with the latter paragraph of the adjacent paragraphs into a single paragraph. That is, the transition information serves as the beginning of the latter paragraph of the adjacent paragraphs.

[0127] When training the large model, user questions from the interaction data are used as input to the large model, and the output of the large model is the response content to the user questions. The parameters of the large model are updated based on the loss of the response content to the user questions output by the large model relative to the sample labels (i.e. the response content to the user questions in the interaction data samples).

[0128] The large model in this application can be a large language model (LLM).

[0129] In an optional embodiment, to maintain the child user's attention and guide their thinking, this application proposes interacting with the child user through a multi-round response + question-guided approach. Based on this, in an optional embodiment, a flowchart illustrating one implementation of the above-mentioned output of child-styled response information generated from voice content is as follows: Figure 4 As shown, it may include:

[0130] Step S401: Determine whether multi-turn interaction is needed for the voice content.

[0131] Optionally, it can be determined whether multiple rounds of interaction are needed for the voice content based on the user's semantics.

[0132] As an example, semantic understanding can be performed on the speech content to obtain a semantic understanding result; this semantic understanding result indicates whether the user needs to perform multiple rounds of interaction with the speech content.

[0133] As an example, semantic understanding can be performed on the speech content to obtain a semantic understanding result. This semantic understanding result indicates whether the user accepts multi-turn interaction with the speech content. If the user accepts multi-turn interaction with the speech content, it is determined that multi-turn interaction with the speech content is necessary; otherwise, it is determined that multi-turn interaction with the speech content is not necessary.

[0134] As an example, semantic understanding can be performed on the voice content and historical interaction content to obtain semantic understanding results; these results indicate whether the user needs to perform multiple rounds of interaction with the voice content.

[0135] As an example, semantic understanding can be performed on the voice content and historical interaction content to obtain a semantic understanding result. This semantic understanding result indicates whether the user accepts multi-turn interaction with the voice content. If the user accepts multi-turn interaction with the voice content, it is determined that multi-turn interaction with the voice content is necessary; otherwise, it is determined that multi-turn interaction with the voice content is not necessary.

[0136] In this embodiment, multi-round interaction will only be performed if the user accepts it; otherwise, multi-round interaction will not be performed.

[0137] Optionally, it can be determined whether multiple rounds of interaction are needed based on the voice content.

[0138] As an example, speech content can be processed using a large model to obtain a processing result; this processing result indicates whether multi-turn interaction is required for the speech content.

[0139] In this embodiment, the need for multiple rounds of interaction is determined based on the voice content itself. Typically, multiple rounds of interaction are performed when there are many responses corresponding to the voice content; otherwise, a single round of interaction is sufficient.

[0140] Optionally, it is possible to determine whether multiple rounds of interaction are needed based on the user's semantics and the voice content.

[0141] As an example, semantic understanding can be performed on the speech content to obtain a semantic understanding result; this result indicates whether the user accepts multi-turn interaction with the speech content. If the semantic understanding result indicates that the user accepts multi-turn interaction with the speech content, the speech content can be processed using a large model to obtain a processing result; this processing result indicates whether multi-turn interaction with the speech content is necessary.

[0142] As an example, semantic understanding can be performed on the speech content and historical interaction content to obtain a semantic understanding result; this result indicates whether the user accepts multi-turn interaction with the speech content. If the semantic understanding result indicates that the user accepts multi-turn interaction with the speech content, the speech content can be processed using a large model to obtain a processing result; this result indicates whether multi-turn interaction with the speech content is necessary.

[0143] In this embodiment, the system first determines whether the user accepts multi-turn interaction with the voice content. Only if the user accepts multi-turn interaction with the voice content will the system further determine whether multi-turn interaction is necessary based on the voice content. If it is determined that multi-turn interaction is necessary based on the voice content, then multi-turn interaction is performed; otherwise, only one-turn interaction is performed, improving the user experience. If the user does not accept multi-turn interaction with the voice content, there is no need to determine whether multi-turn interaction is necessary, and it is directly determined that multi-turn interaction will not be performed.

[0144] Step S402: Based on the judgment result, output the child-styled response information generated based on the voice content.

[0145] In one example, if the judgment result is yes, a child-styled first response and a first question are generated based on the voice content; the first response is part of the target response corresponding to the voice content, and the answer to the first question is related to the remaining content of the target response, thereby guiding the child user to think about the remaining content of the answer to the first question; after outputting the first response and the first question, the process returns to the step of receiving the voice information input by the user and its subsequent steps in order to generate and output the remaining content of the target response; the response content in each round of response information constitutes the target response content.

[0146] If the judgment result is yes, it means that multiple rounds of interaction are needed for the voice content. In this case, the child-styled response information generated based on the voice content includes part of the full content of the response to the voice content (i.e., the target response content), as well as a question about the remaining content of the target response content (i.e., the first question). The first question can guide the child user to think about the remaining content of the target response content, thereby further helping the child user to concentrate.

[0147] In another example, if the determination is yes, a child-styled target response and a second question are generated based on the voice content. The second question guides the child user to think about related content that differs from the target response. After outputting the target response and the second question, the process returns to the step of receiving the user's voice input and its subsequent steps to generate and output related content.

[0148] If the judgment result is yes, it means that multiple rounds of interaction are required based on the voice content. In this case, the child-styled response information generated based on the voice content includes the full content of the response to the voice content (i.e., the target response content), as well as a question related to the target response content (i.e., the first question). The first question can guide the child user to think about content related to the target response content.

[0149] In an optional embodiment, when multiple rounds of interaction are required, if the voice content is a complex question, a child-styled first response and a first question are generated based on the voice content. The first response is a part of the target response corresponding to the voice content, and the answer to the first question is related to the remaining content of the target response. If the voice content is a simple question, a child-styled target response and a second question are generated based on the voice content. The second question is used to guide the child user to think about the relevant content of the voice content, which is different from the target response.

[0150] In other words, when a user asks a complex question, the corresponding response will be lengthy. Outputting all the responses at once might be tedious for children, making it difficult for them to concentrate. Therefore, this application will respond step-by-step, providing a question after each response to guide the child's thinking about the remaining responses. This not only guides the child's thinking but also helps maintain their attention. When a user asks a simple question, the corresponding response is usually shorter. In this case, all responses can be output at once, along with a further question. This question guides the child's thinking about content related to the question, helping them learn new knowledge, expand their knowledge base, and further help maintain their attention.

[0151] If the judgment result is negative, the reply message will be the target reply content.

[0152] If the result is negative, it means that multiple rounds of interaction are not needed for the voice content. In this case, simply output the full response to the voice content.

[0153] Furthermore, the first response content, the first question, the remaining content, the target response content, the second question, and the related content generated based on the voice content satisfy at least one of the following:

[0154] The wording used is childlike;

[0155] The length is less than the preset number of characters;

[0156] The sentence structure uses a simple sentence.

[0157] When the response information generated based on voice content (including the first response content and the first question, the remaining content, the target response content and the second question, related content, etc.) meets at least one of the above conditions, it makes it easier for child users to understand and accept the response information, and makes human-computer interaction smoother.

[0158] In an optional embodiment, the flowchart for generating child-styled response information based on voice content according to the judgment result is as follows: Figure 5 As shown, it may include:

[0159] Step S501: Add the voice content, historical interaction information, and guidance information to the instruction template to obtain the instruction (prompt). The guidance information indicates whether a guided response is needed for the user. The instruction model responds to questions posed by children in a child-friendly style, and also indicates whether a guided response is needed for the user.

[0160] In the first round of interaction, the historical interaction information is empty. After at least two rounds of interaction, the historical interaction information refers to the user input information and the response information output by the human-computer interaction device before the current round.

[0161] The guidance information is derived from the judgment result. If the judgment result is yes, the guidance information indicates that the user needs to be guided to respond, that is, multiple rounds of interaction are required. If the judgment result is no, the guidance information indicates that the user does not need to be guided to respond, that is, multiple rounds of interaction are not required.

[0162] Step S502: Input the instruction into the large model to obtain the child-styled response information generated by the large model.

[0163] In this embodiment, it is not necessary to determine whether the user's question (i.e., the voice content) is a complex question or a simple question. The large model can understand whether the user's question is a simple question or a complex question through fine-tuning the training data, and then output the corresponding response content and question.

[0164] The training dataset used in this application for fine-tuning the large model consists of several interaction data points. Each interaction data point comprises multiple rounds of interaction data. Each round of interaction includes: a user question, a response to the user question, and a guiding question. Specifically, if the user question is simple, the response is the complete answer to the question, and the guiding question is used to guide the user to think about content related to the question (i.e., to guide the user to learn new knowledge). If the user question is complex, the response is a part of the complete answer, and the guiding question is used to guide the user to think about the remaining answer (i.e., to guide the user to think about the remaining answer content).

[0165] When training the large model, user questions from the interaction data are used as input to the large model, and the output of the large model is the response content and guiding question for the user questions. The parameters of the large model are updated based on the loss of the output of the large model (i.e. the response content and guiding question output by the large model for the user questions) relative to the sample labels (i.e. the response content and guiding question output by the large model for the user questions in the interaction data samples).

[0166] The large model in this application can be a large language model (LLM).

[0167] In an optional embodiment, the human-computer interaction method provided in this application may further include:

[0168] Obtain a user's nickname. A user's nickname can be obtained in the following ways:

[0169] The voiceprint of the speech information is extracted and compared with each of the registered voiceprints to determine whether there is a registered voiceprint that matches the voiceprint of the speech information.

[0170] If a registered voiceprint matches the voice information, the nickname corresponding to that registered voiceprint will be used as the user's nickname. When registering a voiceprint, the user can enter their own nickname, thus linking and storing the registered voiceprint with their nickname.

[0171] Accordingly, one way to implement the above-mentioned output of child-styled response information generated from voice content can be:

[0172] Output a child-styled reply message generated based on the voice content and the user's nickname; the reply message includes the user's nickname.

[0173] In this embodiment, the reply information generated by the human-computer interaction device will carry the user's nickname, which increases the sense of intimacy and enhances the user's interactive experience.

[0174] In an optional embodiment, child-styled response information generated based on speech content can be output in a voice manner, specifically:

[0175] The generated child-styled response information can be synthesized into a child-voice-like speech, and then the speech can be output.

[0176] The following example illustrates one implementation of this application:

[0177] In one example, the user inputs the voice message "Why did the apple fall to the ground?"

[0178] Based on the aforementioned voice content, it is determined that the voice content belongs to the encyclopedia Q&A vertical, which is a vertical requiring child-style responses. Based on the voice information, the user is identified as a child user with the nickname "DouDou". Based on the voice content, it is determined that multiple rounds of responses are not required. Therefore, the voice content, historical interaction information, and guidance instructions are added to the prompt template to obtain the prompt, as shown below, which is an example of a prompt provided in this application embodiment:

[0179] "You need to combine historical interaction information to give simple, vivid, easy-to-understand and attractive answers to the questions raised by young children. Now a young child named DouDou has asked you a question, 'Why do apples fall to the ground?' The historical interaction information is {}. Output the answer directly."

[0180] After inputting the above prompt into the large model, the output of the large model is:

[0181] Hi DouDou, apples fall to the ground because of Earth's gravity. Earth is like a big magnet, attracting everything. A scientist named Newton discovered this secret, so now we know why apples fall.

[0182] In another example, the user inputs the audio message "Why can birds fly?"

[0183] Based on the aforementioned voice content, it is determined that the voice content belongs to the encyclopedia Q&A vertical, which is a vertical requiring child-style responses. Based on the voice information, the user is identified as a child user with the nickname "DouDou". Based on the voice content, it is determined that multiple rounds of responses are required. Therefore, the voice content, historical interaction information, and guidance instructions are added to the prompt template to obtain the prompt, as shown below, which is an example of a prompt provided in this application embodiment:

[0184] "You need to provide simple, vivid, easy-to-understand, and engaging answers to questions posed by young children, based on historical interaction information. Now, a young child named DouDou asks you a question: 'Why can birds fly?' The historical interaction information is {}, and the response is guided through multiple rounds." After inputting this prompt into the large model, the model's output is:

[0185] Hi DouDou, birds fly in the blue sky like little boats in the sky because they have a pair of strong wings. When they flap their wings, it's like we paddle with our hands when we swim; this action helps them move upwards and forwards in the air. Do you want to know how birds' wings allow them to fly so high and far?

[0186] To address the above issue, suppose the user inputs the message "Because his wings are very powerful?" via voice.

[0187] The human-computer interaction device determines that the voice content belongs to the encyclopedia Q&A vertical based on the new input voice, which is a vertical requiring a child-style response. Based on the voice information, it identifies the user as a child with the nickname "DouDou." Based on the voice content and historical interaction information, it determines that multiple rounds of responses are needed. Therefore, the human-computer interaction device generates a new prompt, as shown below:

[0188] "You need to provide simple, vivid, easy-to-understand, and engaging answers to questions posed by young children, based on historical interaction information. Now, a young child named DouDou asks you a question: 'Because his wings are very powerful?' The historical interaction information is: {Why can birds fly? Birds fly in the blue sky like small boats in the sky because they have a pair of strong wings. When they flap their wings, it's like we paddle with our hands when swimming; this action helps them move upwards and forwards in the air. Do you want to know how birds' wings allow them to fly so high and far?} Multiple rounds of guided responses." After inputting this prompt into the large model, the output of the large model is:

[0189] That's right, DouDou! Birds have incredibly powerful wings. When a bird flaps its wings, the air creates pressure beneath them, helping it rise and fly forward. It's like when you put your hand out the car window, palm down, the wind pushes your hand up. Now, are you also wondering, besides strong wings, what other secrets allow birds to fly for long periods without getting tired?

[0190] To address the above issue, suppose the user inputs "Is it because he is very light?" via voice.

[0191] The human-computer interaction device determines that the voice content belongs to the encyclopedia Q&A vertical based on the new input voice, which is a vertical requiring a child-style response. Based on the voice information, it identifies the user as a child with the nickname "DouDou." Based on the voice content and historical interaction information, it determines that multiple rounds of responses are not needed. Therefore, the human-computer interaction device generates a new prompt, as shown below:

[0192] "You need to provide simple, vivid, easy-to-understand, and attractive answers to questions posed by young children, based on historical interaction information. Now, a young child named DouDou asks you a question: 'Is it because he is light?' The historical interaction information is: {Why can birds fly? Birds fly in the blue sky like small boats in the sky because they have a pair of strong wings. When they flap their wings, it's like paddling with our hands when we swim; this action helps them rise and move forward in the air. Do you want to know how birds' wings allow them to fly so high and far? Is it because their wings are powerful? That's right! Birds' wings are very powerful. When birds flap their wings, the air creates pressure under their wings, helping them rise and fly forward. It's like when you put your hand out the car window, palm down, the wind pushes your hand up. Now, do you also want to know, besides strong wings, what other secrets allow birds to fly in the air for long periods without getting tired?}, output directly." After inputting this prompt into the large model, the large model outputs the following result:

[0193] That's right, DouDou! A big reason birds are light is because their bones are hollow, making them lighter than other animals and allowing them to fly more easily. This design helps birds fly with effortless grace, almost like dancing. Now you know why birds can fly!

[0194] As mentioned earlier, when interacting with child users, the responses in each round of this application are concise and vivid, with the word count controlled within a certain range, such as around 100 words, making them easy for child users to understand and accept.

[0195] In yet another example, the user inputs the voice message "Why can birds fly?"

[0196] Based on the aforementioned voice content, it is determined that the voice content belongs to the encyclopedia Q&A vertical, which is a vertical requiring child-style responses. Based on the voice information, the user is identified as a child user with the nickname "DouDou". Therefore, the voice content and historical interaction information are added to the prompt template to obtain the prompt, as shown below, which is an example of a prompt provided in this application embodiment:

[0197] "You need to provide simple, vivid, easy-to-understand, and engaging answers to questions posed by young children, based on historical interaction information. Now, a young child named DouDou asks you a question: 'Why can birds fly?' The historical interaction information is {}." After inputting this prompt into the large model, the model's output is:

[0198] Hello DouDou, birds fly in the blue sky like small boats in the sky because they have a pair of strong wings. When they flap their wings, it's like we use our hands to paddle when we swim. This action helps them move upward and forward in the air.

[0199] How do birds' wings allow them to fly so high and so far? It's because birds' wings are very powerful. When a bird flaps its wings, the air creates pressure beneath the wings, helping it rise and fly forward. It's like when you put your hand out the car window, palm down, the wind pushes your hand up.

[0200] Besides their strong wings, what other secret allows birds to fly for long periods without getting tired? It's because birds have hollow bones, making them lighter than other animals and easier to fly in the air. This design helps birds fly with effortless grace, almost like dancing. Now you know why birds can fly.

[0201] In this example, the large model does not engage in multiple rounds of interaction with the user. Instead, it outputs responses in segments. After outputting one segment of responses, before outputting the next segment, it outputs transitional information, such as "How do birds' wings allow them to fly so high and so far?" and "Besides strong wings, what other secrets allow birds to fly in the air for long periods without getting tired?" This tells the user what content will be output next and attracts the attention of child users.

[0202] Corresponding to the method embodiments, this application also provides a human-computer interaction device. A schematic diagram of the structure of the human-computer interaction device provided in the embodiments of this application is shown below. Figure 6 As shown, it may include:

[0203] The system comprises a receiving module 601, a determining module 602, a judging module 603, and an output module 604; wherein,

[0204] The receiving module 601 is used to receive voice information input by the user;

[0205] The determining module 602 is used to determine the vertical domain or topic to which the voice content of the voice information belongs;

[0206] The judgment module 603 is used to determine whether the user is a child user based on the voice information if the vertical domain to which the voice content belongs is a target vertical domain that requires a child-styled response, or if the topic to which the voice content belongs is a target topic that requires a child-styled response.

[0207] The output module 604 is used to output child-styled response information generated based on the voice content if the user is a child user.

[0208] The human-computer interaction device provided in this application embodiment, after receiving voice information input by the user, determines whether the conditions for a child-styled response are met based on the voice information and its voice content. That is, when the voice content input by the user belongs to the target vertical domain or target topic that requires a child-styled response, and the user is a child user, the device replies to the user in a child-style manner. This realizes automatic judgment on whether a child-styled response is needed and improves the intelligence of human-computer interaction.

[0209] In an optional embodiment, when the determining module 602 determines the vertical domain or topic to which the voice content of the voice information belongs, it is used to:

[0210] The speech content is processed by a vertical domain recognition model to obtain the vertical domain to which the speech content belongs;

[0211] Alternatively, the speech content can be processed using a topic recognition model to obtain the topic to which the speech content belongs.

[0212] In an optional embodiment, when the determination module 603 determines whether the user is a child user based on the voice information, it is used to:

[0213] Extract the voiceprint of the speech information;

[0214] If a registered voiceprint that matches the voiceprint exists, obtain the age information corresponding to the registered voiceprint;

[0215] If the age information is within a preset range, the user is determined to be a child user;

[0216] If no registered voiceprint matches the voiceprint, extract the acoustic features of the speech information;

[0217] The acoustic features are input into the growth stage recognition model to obtain the recognition result output by the growth stage recognition model, and the recognition result indicates whether the user is a child.

[0218] In an optional embodiment, when the determination module 603 determines whether the user is a child user based on the voice information, it is used to:

[0219] Extract the acoustic features of the speech information;

[0220] The acoustic features are input into the growth stage recognition model to obtain the recognition result output by the growth stage recognition model, and the recognition result indicates whether the user is a child.

[0221] In an optional embodiment, when the output module 604 outputs child-styled response information generated based on the speech content, it is used to:

[0222] Based on the voice content, one or more child-styled response segments are generated; wherein, when multiple child-styled response segments are generated, adjacent child-styled response segments are connected by transition information, wherein the transition information is a summary of the latter child-styled response segment in the two adjacent child-styled response segments.

[0223] Output one or more paragraphs of child-styled response content.

[0224] In an optional embodiment, each child-styled response satisfies at least one of the following:

[0225] The wording should be childlike;

[0226] The length is less than the preset number of characters;

[0227] The sentence structure uses a simple sentence.

[0228] In an optional embodiment, when the output module 604 outputs child-styled response information generated based on the speech content, it is used to:

[0229] Determine whether multiple rounds of interaction are needed for the voice content;

[0230] Based on the judgment result, output a child-styled response message generated from the spoken content; specifically:

[0231] If the determination result is yes, a child-styled first response content and a first question are generated based on the voice content. The first response content is a part of the target response content corresponding to the voice content, and the answer to the first question is related to the remaining content of the target response content. After outputting the first response content and the first question, the process returns to the step of receiving voice information input by the user and its subsequent steps to generate and output the remaining content. Alternatively, a child-styled target response content and a second question are generated based on the voice content. The second question is used to guide the child user to think about the related content of the voice content, and the related content is different from the target response content. After outputting the target response content and the second question, the process returns to the step of receiving voice information input by the user and its subsequent steps to generate and output the related content.

[0232] If the judgment result is negative, generate the target response content in a child-style based on the voice content, and output the target response content.

[0233] In an optional embodiment, the first response content and the first question generated based on the voice content, the remaining content, the target response content and the second question, and the related content satisfy at least one of the following:

[0234] The wording used is childlike;

[0235] The length is less than the preset number of characters;

[0236] The sentence structure uses a simple sentence.

[0237] In an optional embodiment, when the output module 604 determines whether multiple rounds of interaction are needed for the speech content, it is used to:

[0238] The speech content is subjected to semantic understanding to obtain a semantic understanding result; the semantic understanding result indicates whether multiple rounds of response are required for the speech content.

[0239] Alternatively, semantic understanding can be performed on the speech content and historical interaction content to obtain a semantic understanding result; this semantic understanding result indicates whether multiple rounds of responses are needed for the speech content.

[0240] Alternatively, semantic understanding can be performed on the speech content, or semantic understanding can be performed on the speech content and historical interaction content; if the semantic understanding result indicates that the user accepts multiple rounds of responses to the speech content, the speech content is processed through a large model to obtain a processing result; the processing result indicates whether multiple rounds of responses to the speech content are required.

[0241] In an optional embodiment, when the output module 604 outputs child-styled response information generated based on the speech content, it is used to:

[0242] Output a child-styled reply message generated based on the voice content and the user's nickname; the reply message carries the user's nickname.

[0243] This application also provides an electronic device in its embodiments. (See reference...) Figure 7 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals 7 such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0244] like Figure 7 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. When the electronic device is powered on, the RAM 703 also stores various programs and data required for the operation of the electronic device. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0245] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, memory cards, hard drives, etc.; and communication devices 709. Communication device 709 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0246] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the human-computer interaction methods provided in this application.

[0247] This application also provides a computer-readable storage medium that carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the human-computer interaction methods provided in this application.

[0248] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0249] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0250] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0251] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A human-computer interaction method, characterized in that, include: Receive voice information input by the user; Determine the vertical domain or topic to which the voice content of the voice information belongs; If the vertical domain to which the voice content belongs is the target vertical domain that requires a child-styled response, or if the topic to which the voice content belongs is the target topic that requires a child-styled response, determine whether the user is a child user based on the voice information; If the user is a child, output a child-styled response message generated based on the voice content.

2. The method according to claim 1, characterized in that, Determining the vertical domain or topic to which the voice content of the voice information belongs includes: The speech content is processed by a vertical domain recognition model to obtain the vertical domain to which the speech content belongs; Alternatively, the speech content can be processed using a topic recognition model to obtain the topic to which the speech content belongs.

3. The method according to claim 1, characterized in that, The step of determining whether the user is a child user based on the voice information includes: Extract the voiceprint of the speech information; If a registered voiceprint that matches the voiceprint exists, obtain the age information corresponding to the registered voiceprint; If the age information is within a preset range, the user is determined to be a child user; If no registered voiceprint matches the voiceprint, extract the acoustic features of the speech information; The acoustic features are input into the growth stage recognition model to obtain the recognition result output by the growth stage recognition model, and the recognition result indicates whether the user is a child.

4. The method according to claim 1, characterized in that, The step of determining whether the user is a child user based on the voice information includes: Extract the acoustic features of the speech information; The acoustic features are input into the growth stage recognition model to obtain the recognition result output by the growth stage recognition model, and the recognition result indicates whether the user is a child.

5. The method according to claim 1, characterized in that, The output, based on the voice content, generates child-styled response information, including: Based on the voice content, one or more child-styled response segments are generated; wherein, when multiple child-styled response segments are generated, adjacent child-styled response segments are connected by transition information, wherein the transition information is a summary of the latter child-styled response segment in the two adjacent child-styled response segments. Output one or more paragraphs of child-styled response content.

6. The method according to claim 5, characterized in that, Each child-styled response must meet at least one of the following criteria: The wording should be childlike; The length is less than the preset number of characters; The sentence structure uses a simple sentence.

7. The method according to claim 1, characterized in that, The output, based on the voice content, generates child-styled response information, including: Determine whether multiple rounds of interaction are needed for the voice content; Based on the judgment result, output a child-styled response message generated from the spoken content; specifically: If the determination result is yes, a child-styled first response content and a first question are generated based on the voice content. The first response content is a part of the target response content corresponding to the voice content, and the answer to the first question is related to the remaining content of the target response content. After outputting the first response content and the first question, the process returns to the step of receiving voice information input by the user and its subsequent steps to generate and output the remaining content. Alternatively, a child-styled target response content and a second question are generated based on the voice content. The second question is used to guide the child user to think about the related content of the voice content, and the related content is different from the target response content. After outputting the target response content and the second question, the process returns to the step of receiving voice information input by the user and its subsequent steps to generate and output the related content. If the judgment result is negative, generate the target response content in a child-style based on the voice content, and output the target response content.

8. The method according to claim 7, characterized in that, The first response content and the first question generated based on the voice content, the remaining content, the target response content and the second question, and the related content satisfy at least one of the following: The wording used is childlike; The length is less than the preset number of characters; The sentence structure uses a simple sentence.

9. The method according to claim 7, characterized in that, The determination of whether multiple rounds of interaction are needed for the voice content includes: The speech content is subjected to semantic understanding to obtain a semantic understanding result; the semantic understanding result indicates whether multi-turn interaction is required for the speech content. Alternatively, semantic understanding can be performed on the speech content and historical interaction content to obtain a semantic understanding result; this semantic understanding result indicates whether multiple rounds of interaction are required for the speech content. Alternatively, semantic understanding can be performed on the speech content, or semantic understanding can be performed on the speech content and historical interaction content; if the semantic understanding result indicates that the user accepts multi-turn interaction for the speech content, the speech content is processed to obtain a processing result; the processing result indicates whether multi-turn interaction for the speech content is required.

10. The method according to any one of claims 1-9, characterized in that, The output, based on the voice content, generates child-styled response information, including: Output a child-styled reply message generated based on the voice content and the user's nickname; the reply message carries the user's nickname.

11. A human-computer interaction device, characterized in that, include: The receiving module is used to receive voice information input by the user; The determination module is used to determine the vertical domain or topic to which the voice content of the voice information belongs; The judgment module is used to determine whether the user is a child user based on the voice information if the vertical domain to which the voice content belongs is a target vertical domain that requires a child-styled response, or if the topic to which the voice content belongs is a target topic that requires a child-styled response. The output module is used to output child-styled response information generated based on the voice content if the user is a child user.

12. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the human-computer interaction method as described in any one of claims 1 to 10.

13. A computer program product, characterized in that, It includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the human-computer interaction method as described in any one of claims 1 to 10.

14. A computer storage medium, characterized in that, The computer storage medium carries one or more computer programs, which, when executed by an electronic device, enable the electronic device to implement the human-computer interaction method as described in any one of claims 1 to 10.