Voice interaction method for security scenario, apparatus, electronic device and storage medium

By acquiring security video to determine the identity and behavior type of visitors, generating and playing matching voice interaction strategies, the security issues of visitors in security scenarios are solved, and security and deterrence are improved.

WO2026114296A1PCT designated stage Publication Date: 2026-06-04SHENZHEN OCEANWING SMART INNOVATIONS TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SHENZHEN OCEANWING SMART INNOVATIONS TECHNOLOGY CO LTD
Filing Date
2025-11-27
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

In security scenarios, visitors cannot be separated from the real-time participation of the person being visited, which may affect property safety or personal safety.

Method used

By acquiring security video, the system determines the identity and behavior type of visitors, generates a matching voice interaction strategy, and controls the security agent to play interactive voice messages.

Benefits of technology

It improves security during visits in security scenarios by generating voice interaction strategies that match the visitors, thereby enhancing deterrence and security for visitors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025138003_04062026_PF_FP_ABST
    Figure CN2025138003_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A voice interaction method for a security scenario, an apparatus, an electronic device and a storage medium, the method comprising: acquiring a first security video containing an image of a first visiting subject (101); on the basis of the first security video, determining a first identity type and a first behavior type of the first visiting subject (102); determining a first voice interaction policy matching the first visiting subject (103); on the basis of the first voice interaction policy, the first identity type and the first behavior type, generating a first interaction voice for interaction (104); and controlling a security agent to play back the first interaction voice, wherein the security agent may be a smart doorbell (105). Thus, the interaction voice for interaction may be generated on the basis of the voice interaction policy matching the visiting subject, the identity type and the behavior type of the visiting subject, and then the security agent is controlled to play back the interaction voice. In this way, the security during subject visiting may be improved in security scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Voice interaction methods, devices, electronic equipment and storage media in security scenarios

[0001] This application claims priority to Chinese Patent Application No. 202411759066.9, filed on November 29, 2024, entitled "Voice Interaction Method, Apparatus, Electronic Device and Storage Medium in Security Scenarios", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of security, and more particularly to a voice interaction method, device, electronic device, and storage medium in a security scenario. Background Technology

[0003] In the security field, when visitors arrive, the interviewee needs to have a face-to-face conversation with the visitor, or conduct a real-time conversation using the voice transmission function of devices such as microphones and speakers.

[0004] However, the above methods cannot be conducted without the real-time participation of the interviewees. If the interviewee does not respond after the visitor arrives, the visitor may realize that the interviewee is not present, which could affect the interviewee's property safety or the personal safety of vulnerable groups related to the interviewee.

[0005] It is evident that improving the security of visitors in security scenarios is a technical issue worthy of attention. Summary of the Invention

[0006] This application provides a voice interaction method, device, electronic device, and storage medium for security scenarios, in order to improve the security of visitors in security scenarios.

[0007] In a first aspect, embodiments of this application provide a voice interaction method in a security scenario, the method comprising:

[0008] Acquire the first security video containing images of the first visitor;

[0009] Based on the first security video, determine the first identity type and first behavior type of the first visitor;

[0010] Determine a first voice interaction strategy that matches the first visitor;

[0011] According to the first voice interaction strategy, a first interactive voice for interaction is generated based on the first identity type and the first behavior type;

[0012] Control the security intelligent agent to play the first interactive voice.

[0013] In conjunction with the first aspect, in the first possible implementation of the first aspect, determining the first identity type of the first visitor based on the first security video includes:

[0014] The first security video is input into a pre-trained first model to obtain the first identity type of the first visitor.

[0015] In conjunction with the first aspect, in the second possible implementation of the first aspect, determining the first identity type of the first visitor based on the first security video includes:

[0016] Obtain images of objects of different identity types to obtain the first image set;

[0017] Extract the feature data of the first security video;

[0018] Calculate the similarity between the feature data and the feature data of each image in the first image set;

[0019] The identity type corresponding to the highest similarity value among the obtained similarities is determined as the first identity type of the first visitor.

[0020] In conjunction with the first aspect, in the third possible implementation of the first aspect, based on the first security video, determining the first behavior type of the first visitor includes:

[0021] The first security video is input into a pre-trained second model to obtain the first behavior type of the first visitor.

[0022] In conjunction with the first aspect, in the fourth possible implementation of the first aspect, based on the first security video, determining the first behavior type of the first visitor includes:

[0023] Obtain images of objects with different behavior types to obtain a second image set;

[0024] Extract the feature data of the first security video;

[0025] Calculate the similarity between the feature data and the feature data of each image in the second image set;

[0026] The behavior type corresponding to the highest similarity value among the obtained similarities is determined as the first behavior type of the first visitor.

[0027] In conjunction with the first aspect, in the fifth possible implementation of the first aspect, the first voice interaction strategy includes the pronunciation style of a single speaker in voice interaction; and

[0028] The step of generating a first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy includes:

[0029] Based on the first identity type and the first behavior type, a first interactive voice with the pronunciation style is generated for interaction.

[0030] In conjunction with the first aspect, in the sixth possible implementation of the first aspect, generating the first interactive voice for interaction and having the pronunciation style based on the first identity type and the first behavior type includes:

[0031] Obtain the feature data of the pronunciation style;

[0032] The feature data of the pronunciation style, the first identity type, and the first behavior type are input into a pre-trained fourth model to obtain a first interactive voice with the pronunciation style for interaction.

[0033] In conjunction with the first aspect, in the seventh possible implementation of the first aspect, generating the first interactive voice for interaction and having the pronunciation style based on the first identity type and the first behavior type includes:

[0034] Obtain the fifth model corresponding to the pronunciation style;

[0035] The first identity type and the first behavior type are input into the fifth model corresponding to the pronunciation style to obtain a first interactive voice that is used for interaction and has the pronunciation style.

[0036] In conjunction with the first aspect, in the eighth possible implementation of the first aspect, the first voice interaction strategy includes the voice timbre for performing the voice interaction; and

[0037] The step of generating a first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy includes:

[0038] Based on the first identity type and the first behavior type, a first interactive voice with the aforementioned voice timbre is generated for interaction.

[0039] In conjunction with the first aspect, in the ninth possible implementation of the first aspect, generating a first interactive voice for interaction and having the pronunciation timbre based on the first identity type and the first behavior type includes:

[0040] Obtain the characteristic data of the pronunciation timbre;

[0041] The feature data of the pronunciation timbre, the first identity type, and the first behavior type are input into the pre-trained sixth model to obtain the first interactive voice with the pronunciation timbre for interaction.

[0042] In conjunction with the first aspect, in the tenth possible implementation of the first aspect, generating a first interactive voice for interaction and having the pronunciation timbre based on the first identity type and the first behavior type includes:

[0043] Obtain the seventh model corresponding to the pronunciation timbre;

[0044] The first identity type and the first behavior type are input into the seventh model corresponding to the voice timbre to obtain a first interactive voice that is used for interaction and has the pronunciation timbre.

[0045] In conjunction with the first aspect, in an eleventh possible implementation of the first aspect, before determining the first voice interaction strategy matching the first visitor, the method further includes:

[0046] Determine the interviewee of the first visitor;

[0047] The object information of the interviewees is determined, wherein the object information includes at least one of the following: the interviewee's voice timbre, the interviewee's preference information, the interviewee's age, the interviewee's gender, and the number of interviewees; and

[0048] The first voice interaction strategy for determining the match with the first visitor includes:

[0049] Based on the object information, a first voice interaction strategy matching the first visiting object is determined.

[0050] In conjunction with the first aspect, in the twelfth possible implementation of the first aspect, when the object information includes the preference information of the interviewee, determining the first voice interaction strategy matching the first interviewee based on the object information includes:

[0051] Determine whether the preference information includes a voice interaction strategy triggered by the first visitor;

[0052] If the preference information includes a voice interaction strategy triggered by the first visitor, the voice interaction strategy triggered by the first visitor is determined as a first voice interaction strategy that matches the first visitor.

[0053] If the preference information does not include the voice interaction strategy triggered by the first visitor, a first voice interaction strategy matching the first visitor is generated based on the preference information.

[0054] In conjunction with the first aspect, in a thirteenth possible implementation of the first aspect, before determining the first voice interaction strategy matching the first visitor, the method further includes:

[0055] Based on the first security video, determine the environmental information of the environment where the first visitor is located; and

[0056] The first voice interaction strategy for determining the match with the first visitor includes:

[0057] Based on at least one of the environmental information, the first identity type, and the first behavior type, a first voice interaction strategy matching the first visitor is determined.

[0058] In conjunction with the first aspect, in the fourteenth possible implementation of the first aspect, determining the first identity type and first behavior type of the first visitor based on the first security video includes:

[0059] The first security video is input into a multimodal big data model to determine the first identity type and first behavior type of the first visitor; and / or

[0060] The first voice interaction strategy for determining the match with the first visitor includes:

[0061] Using a multimodal large model, determine a first voice interaction strategy that matches the first visitor; and / or

[0062] The step of generating a first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy includes:

[0063] The first voice interaction strategy, the first identity type, and the first behavior type are input into the multimodal large model to generate the first interactive voice for interaction.

[0064] In conjunction with the first aspect, in the fifteenth possible implementation of the first aspect, after the control security agent plays the first interactive voice, the method further includes:

[0065] Obtain the recorded video of the security intelligent agent playing the first interactive voice;

[0066] Determine the feedback information from the preset object to the recorded video;

[0067] Acquire a second security video containing images of a second visitor;

[0068] Based on the second security video, determine the second identity type and second behavior type of the second visitor;

[0069] Based on the feedback information, a second voice interaction strategy matching the second visitor is determined;

[0070] According to the second voice interaction strategy, based on the second identity type and the second behavior type, a second interactive voice is generated for interacting with the second visitor.

[0071] Control the security intelligent agent to play the second interactive voice.

[0072] In conjunction with the first aspect, in the sixteenth possible implementation of the first aspect, the security intelligent agent is a smart doorbell.

[0073] Secondly, embodiments of this application provide a voice interaction device for security scenarios, the device comprising:

[0074] The first acquisition unit is used to acquire a first security video containing an image of the first visitor.

[0075] The first determining unit is used to determine the first identity type and the first behavior type of the first visitor based on the first security video.

[0076] The second determining unit is used to determine a first voice interaction strategy that matches the first visitor.

[0077] The first generation unit is configured to generate a first interactive voice for interaction based on the first identity type and the first behavior type, according to the first voice interaction strategy.

[0078] The first control unit is used to control the security intelligent agent to play the first interactive voice.

[0079] In one possible implementation, the first voice interaction strategy includes the pronunciation style of a single speaker during voice interaction; and

[0080] The step of generating a first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy includes:

[0081] Based on the first identity type and the first behavior type, a first interactive voice with the pronunciation style is generated for interaction.

[0082] Thirdly, embodiments of this application provide an electronic device, including:

[0083] Memory, used to store computer programs;

[0084] A processor is configured to execute a computer program stored in the memory, and when the computer program is executed, to implement any embodiment of the voice interaction method in the security scenario described in the first aspect of this application.

[0085] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the method of any embodiment of the voice interaction method in the security scenario described in the first aspect above.

[0086] The voice interaction method in a security scenario provided in this application embodiment can acquire a first security video containing an image of a first visitor. Then, based on the first security video, a first identity type and a first behavior type of the first visitor are determined. Next, a first voice interaction strategy matching the first visitor is determined. Subsequently, according to the first voice interaction strategy, a first interactive voice for interaction is generated based on the first identity type and the first behavior type. Finally, the security intelligent agent is controlled to play the first interactive voice. Therefore, according to the voice interaction strategy matching the visitor, and based on the visitor's identity type and behavior type, interactive voice for interaction can be generated, and the security intelligent agent can be controlled to play the interactive voice, thus improving the security of visitors in a security scenario. Attached Figure Description

[0087] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0088] Figure 1 is a flowchart illustrating a voice interaction method in a security scenario provided in an embodiment of this application;

[0089] Figure 2 is a flowchart of a method for determining a first identity type according to an embodiment of this application;

[0090] Figure 3 is a flowchart of another method for determining a first identity type provided in an embodiment of this application;

[0091] Figure 4 is a flowchart of a method for determining a first behavior type according to an embodiment of this application;

[0092] Figure 5 is a flowchart of another method for determining the first behavior type provided in an embodiment of this application;

[0093] Figure 6 is a flowchart of a method for generating a first interactive voice with a pronunciation style according to an embodiment of this application;

[0094] Figure 7 is a flowchart of another method for generating first interactive speech with pronunciation style provided in an embodiment of this application;

[0095] Figure 8 is a flowchart of another method for generating a first interactive voice with a pronunciation style provided in an embodiment of this application;

[0096] Figure 9 is a flowchart of a method for generating a first interactive voice with a vocal timbre according to an embodiment of this application;

[0097] Figure 10 is a flowchart of another method for generating a first interactive voice with a vocal timbre provided in an embodiment of this application;

[0098] Figure 11 is a flowchart of another method for generating a first interactive voice with a vocal timbre provided in an embodiment of this application;

[0099] Figure 12 is a flowchart of a method for determining a first voice interaction strategy provided in an embodiment of this application;

[0100] Figure 13 is a flowchart of another embodiment of the voice interaction method in a security scenario provided by the present application;

[0101] Figure 14 is a flowchart of another embodiment of the voice interaction method in a security scenario provided by the present application;

[0102] Figure 15 is a flowchart illustrating another voice interaction method in a security scenario provided in the embodiments of this application;

[0103] Figure 16 is a flowchart of a method for generating a first voice interaction strategy according to an embodiment of this application;

[0104] Figure 17A is a schematic diagram of an application scenario of a voice interaction method in a security scenario provided by an embodiment of this application;

[0105] Figure 17B is a schematic diagram of another application scenario of the voice interaction method in a security scenario provided by the embodiments of this application;

[0106] Figure 18 is a structural schematic diagram of a voice interaction device in a security scenario provided in an embodiment of this application;

[0107] Figure 19 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0108] Various exemplary embodiments of this application will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this application.

[0109] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of this application are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor do they indicate the logical order between them.

[0110] It should also be understood that in this embodiment, "multiple" can refer to two or more, and "at least one" can refer to one, two or more.

[0111] It should also be understood that any component, data or structure mentioned in the embodiments of this application can generally be understood as one or more unless explicitly defined or given contrary guidance in the context.

[0112] Furthermore, the term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this application generally indicates that the preceding and following related objects have an "or" relationship.

[0113] It should also be understood that the description of the various embodiments in this application emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.

[0114] The following description of at least one exemplary embodiment is merely illustrative and is not intended to limit the scope of this application or its application or use.

[0115] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0116] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0117] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. To facilitate understanding of the embodiments of this application, the application will be described in detail below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0118] To address the technical problem of improving security when visitors arrive in security scenarios, this application provides a voice interaction method, device, electronic device, and storage medium for security scenarios, which can improve security when visitors arrive in security scenarios.

[0119] Figure 1 is a flowchart illustrating a voice interaction method in a security scenario according to an embodiment of this application. This method can be applied to one or more electronic devices such as smart doorbells, smartphones, smart speakers, laptops, desktop computers, portable computers, and servers. Furthermore, the execution entity of this method can be hardware or software. When the execution entity is hardware, it can be one or more of the aforementioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the execution entity is software, this method can be implemented as multiple software programs or software modules, or as a single software program or software module. Moreover, this method can be executed in conjunction with software and hardware, without specific limitations herein.

[0120] As shown in Figure 1, the method specifically includes:

[0121] Step 101: Obtain the first security video containing the image of the first visitor.

[0122] In this embodiment, the first visitor can be any one or more visitors. Visitors can be delivery personnel, police officers, neighbors, or other individuals; they can also be mobile electronic devices such as drones or vehicles; or even pets.

[0123] The first security video can be a video containing the image of the first visitor. This first security video can be acquired via a camera. For example, such cameras can be installed at building entrances, corridors, elevator lobbies, stairwells, etc. Alternatively, the camera can be installed on a door, such as a door lock or doorbell.

[0124] Step 102: Based on the first security video, determine the first identity type and the first behavior type of the first visitor.

[0125] In this embodiment, the first identity type can represent the identity of the first visitor. For example, if the first visitor is a neighbor, the first identity type can represent a neighbor; if the first visitor is a courier, the first identity type can represent a courier; if the first visitor is the homeowner, the first identity type can represent the homeowner; and if the identity of the first visitor is not identified, the first identity type can represent a stranger.

[0126] The first behavior type can represent the behavior of the first visitor. For example, if the first visitor's behavior is delivering a package, the first behavior type can represent delivering a package; if the first visitor's behavior is loitering, the first behavior type can represent loitering.

[0127] Here, multiple methods can be used to determine the first identity type of the first visitor based on the first security video.

[0128] As an example, the first identity type of the first visitor can be determined using the process shown in Figure 2. Referring to Figure 2, a flowchart of a method for determining a first identity type according to an embodiment of this application is provided. As shown in Figure 2, the process may include the following steps:

[0129] Step 201: Input the first security video into the pre-trained first model to obtain the first identity type of the first visitor.

[0130] In this step, the first security video can be input into a pre-trained first model to determine the first identity type of the first visitor. The first model represents the correspondence between the security video and the identity type of the visitor represented by the images in the security video. The first model can be an artificial intelligence model trained using multiple first training samples, such as a Large Language Model (LLM). The first training samples can include the security video and the identity type of the visitor represented by the images in the security video.

[0131] As another example, the first identity type of the first visitor can be determined through the process shown in Figure 3. Referring to Figure 3, a flowchart of another method for determining the first identity type provided in an embodiment of this application is shown. As shown in Figure 3, the process may include the following steps:

[0132] Step 301: Obtain images of objects of different identity types to obtain the first image set.

[0133] Step 302: Extract the feature data of the first security video.

[0134] Step 303: Calculate the similarity between the above feature data and the feature data of each image in the first image set.

[0135] Step 304: Determine the identity type corresponding to the highest similarity value among the obtained similarity scores as the first identity type of the first visitor.

[0136] The following provides a unified explanation of steps 301 to 304:

[0137] In this embodiment, images of objects with different identity types can be acquired first to obtain a first image set. The first image set may include multiple acquired images. Each image in the first image set may correspond to one identity type. Next, feature data from a first security video can be extracted, and the similarity between this feature data and the feature data of each image in the first image set can be calculated. Then, the identity type corresponding to the highest similarity value among the calculated similarities is determined as the first identity type of the first visiting object.

[0138] The feature data in the above examples may include, but is not limited to, facial features, iris features, fingerprint features, gait features, etc.

[0139] In addition, the similarity in the above examples can be obtained by calculating Euclidean distance, cosine similarity, and Hamming distance.

[0140] In addition, various methods can be used to determine the first behavior type of the first visitor based on the first security video.

[0141] As an example, the first behavior type of the first visitor can be determined using the flowchart shown in Figure 4. Referring to Figure 4, a flowchart of a method for determining the first behavior type provided in an embodiment of this application is shown. As shown in Figure 4, the process may include the following steps:

[0142] Step 401: Input the first security video into the pre-trained second model to obtain the first behavior type of the first visitor.

[0143] In this step, the first security video can be input into a pre-trained second model to determine the first behavior type of the first visitor. The second model represents the correspondence between the security video and the behavior type of the visitor represented by the images in the security video. The second model can be an artificial intelligence model trained using multiple second training samples, such as a large language model. The second training samples can include the security video and the behavior types of the visitor represented by the images in the security video.

[0144] As another example, the first behavior type of the first visitor can be determined using the flowchart shown in Figure 5. Referring to Figure 5, a flowchart of another method for determining the first behavior type provided in an embodiment of this application is shown. As shown in Figure 5, the process may include the following steps:

[0145] Step 501: Obtain images of objects with different behavior types to obtain a second image set.

[0146] Step 502: Extract the feature data of the first security video.

[0147] Step 503: Calculate the similarity between the above feature data and the feature data of each image in the second image set.

[0148] Step 504: Determine the behavior type corresponding to the highest similarity value among the obtained similarity scores as the first identity type of the first visitor.

[0149] The following provides a unified explanation of steps 501 to 504:

[0150] In this embodiment, images of objects with different behavior types can be acquired first to obtain a second image set. The second image set may include multiple acquired images, and each image in the second image set may correspond to a behavior type. Next, feature data from the first security video can be extracted, and the similarity between this feature data and the feature data of each image in the second image set can be calculated. Then, the behavior type corresponding to the highest similarity value among the calculated similarities is determined as the first behavior type of the first visiting object.

[0151] The feature data in the above examples may include, but is not limited to, joint features, limb features, etc.

[0152] In addition, the similarity in the above examples can be obtained by calculating Euclidean distance, cosine similarity, and Hamming distance.

[0153] In some cases, steps one and two can be executed in parallel, or step one can be executed first, and step two can be executed only if the first identity type determined in step one is a preset identity type. The preset identity type can represent a stranger.

[0154] The first step is to determine the first identity type of the first visitor based on the first security video. The second step is to determine the first behavior type of the first visitor based on the first security video.

[0155] Step 103: Determine the first voice interaction strategy that matches the first visitor.

[0156] In this embodiment, the first voice interaction strategy may be a voice interaction strategy that matches the first visitor.

[0157] The voice interaction strategy (including the first voice interaction strategy) may include at least one of the following: the pronunciation style of the speech object (e.g., polite, humorous, stern, friendly) and the timbre of the speech object (e.g., the timbre of an adult male, the timbre of a child, etc.).

[0158] Here, various methods can be used to determine the first voice interaction strategy that matches the first visitor. Please refer to the description below for details, which will not be elaborated upon here.

[0159] Step 104: In accordance with the first voice interaction strategy, generate the first interactive voice for interaction based on the first identity type and the first behavior type.

[0160] In this embodiment, the first interactive voice can conform to the aforementioned first voice interaction strategy. For example, if the first voice interaction strategy indicates that the pronunciation style is polite, the first interactive voice can be polite, such as "Hello, please leave the package at the door." If the first voice interaction strategy indicates that the pronunciation style is an adult male's tone, the first interactive voice can have that adult male's tone, such as saying "Leave as soon as possible" in a stern adult male's tone.

[0161] Here, multiple methods can be used to generate the first interactive voice for interaction based on the first identity type and the first behavior type, according to the first voice interaction strategy.

[0162] As an example, a first voice interaction strategy, a first identity type, and a first behavior type can be input into a pre-trained third model to generate a first interactive voice for interaction.

[0163] The third model can represent the correspondence between voice interaction strategies, identity types, and behavior types. For example, the third model can be an artificial intelligence model trained using multiple third training samples, such as a large language model. The third training samples can include voice interaction strategies, identity types, and behavior types.

[0164] Step 105: Control the security smart agent to play the first interactive voice. The security smart agent can be a smart doorbell.

[0165] In this embodiment, after generating the first interactive voice, the security intelligent agent can be further controlled to play the first interactive voice.

[0166] Among them, the security intelligent agent can be any one or more intelligent devices with voice playback function.

[0167] As an example, security smart agents can be: smart doors, smart door locks, smart speakers, etc.

[0168] In some optional implementations of this embodiment, the security smart agent is a smart doorbell.

[0169] It is understandable that, among the above optional implementation methods, the first interactive voice can be played through a smart doorbell to improve the security of visitors in security scenarios.

[0170] In some optional implementations of this embodiment, the first voice interaction strategy includes the pronunciation style of a single speaking object for voice interaction.

[0171] Based on this, the process shown in Figure 6 can be used to generate a first interactive voice with the aforementioned pronunciation style. Referring to Figure 6, a flowchart of a method for generating a first interactive voice with a pronunciation style is provided in an embodiment of this application. As shown in Figure 6, the process may include the following steps:

[0172] Step 601: Based on the first identity type and the first behavior type mentioned above, generate a first interactive voice for interaction that has the above pronunciation style.

[0173] In this step, the first interactive voice for interaction can be generated according to the first voice interaction strategy, based on the first identity type and the first behavior type: the first interactive voice for interaction with the pronunciation style is generated based on the first identity type and the first behavior type.

[0174] As an example, the process shown in Figure 7 can be used to generate a first interactive voice with the aforementioned pronunciation style. Referring to Figure 7, a flowchart of another method for generating a first interactive voice with a pronunciation style is provided in an embodiment of this application. As shown in Figure 7, the process may include the following steps:

[0175] Step 701: Obtain feature data of pronunciation style.

[0176] Step 702: Input the feature data of the above pronunciation style, the above first identity type and the first behavior type into the pre-trained fourth model to obtain the first interactive speech with the above pronunciation style for interaction.

[0177] The following provides a unified explanation of steps 701 and 702:

[0178] In this embodiment of the application, the feature data of the pronunciation style, the first identity type and the first behavior type can be input into a pre-trained fourth model to generate a first interactive voice with the pronunciation style for interaction.

[0179] The fourth model can represent the correspondence between pronunciation style, identity type, and behavior type. For example, the fourth model can be an artificial intelligence model trained using multiple fourth training samples, such as a large language model.

[0180] The fourth training sample may include feature data on pronunciation style, identity type, and behavior type.

[0181] As another example, a first interactive voice with the aforementioned pronunciation style can be generated through the process shown in Figure 8. Referring to Figure 8, a flowchart of another method for generating a first interactive voice with a pronunciation style provided in an embodiment of this application is shown. As shown in Figure 8, the process may include the following steps:

[0182] Step 801: Obtain the fifth model corresponding to the above pronunciation style.

[0183] Step 802: Input the first identity type and the first behavior type into the fifth model corresponding to the pronunciation style to obtain the first interactive voice that is used for interaction and has the pronunciation style.

[0184] The following provides a unified explanation of steps 801 and 802:

[0185] In this embodiment, different pronunciation styles can correspond to different fifth models. Therefore, the first identity type and the first behavior type can be input into the fifth model corresponding to the pronunciation style to generate a first interactive voice with the pronunciation style for interaction.

[0186] The fifth model can represent the correspondence between identity types and behavior types. For example, the fifth model can be an artificial intelligence model trained using training samples that include identity types and behavior types, such as a large language model.

[0187] For example, if the first identity type is "suspicious person" and the first behavior type is "loitering," then the first interactive voice could be a menacing tone (corresponding to the pronunciation style described above) used to drive someone away. This helps ensure the safety of family members, especially when only children, the elderly, or women are present at home.

[0188] It is understandable that, among the above optional implementation methods, different pronunciation styles of first interactive voice can be generated for different first visitors, and the security intelligent agent can play the first interactive voice with different pronunciation styles. In this way, the matching between the first interactive voice and the first visitor can be improved, thereby further improving the security of the visitor in the security scenario.

[0189] In some optional implementations of this embodiment, the first voice interaction strategy includes the voice timbre for performing voice interaction.

[0190] Based on this, a first interactive voice with the aforementioned pronunciation timbre can be generated using the process shown in Figure 9. Referring to Figure 9, a flowchart of a method for generating a first interactive voice with pronunciation timbre is provided in an embodiment of this application. As shown in Figure 9, the process may include the following steps:

[0191] Step 901: Based on the first identity type and the first behavior type mentioned above, generate a first interactive voice for interaction that has the aforementioned pronunciation timbre.

[0192] In this step, the first interactive voice for interaction can be generated according to the first voice interaction strategy, based on the first identity type and the first behavior type, in the following manner: Based on the first identity type and the first behavior type, a first interactive voice for interaction with the aforementioned voice timbre can be generated.

[0193] As an example, a first interactive voice with the aforementioned pronunciation timbre can be generated through the process shown in Figure 10. Referring to Figure 10, a flowchart of another method for generating a first interactive voice with a pronunciation timbre is provided in an embodiment of this application. As shown in Figure 10, the process may include the following steps:

[0194] Step 1001: Obtain the characteristic data of the above-mentioned pronunciation timbre.

[0195] Step 1002: Input the aforementioned pronunciation timbre feature data, the aforementioned first identity type, and the aforementioned first behavior type into the pre-trained sixth model to obtain the first interactive speech for interaction that has the aforementioned pronunciation timbre.

[0196] The following provides a unified explanation of steps 1001 and 1002:

[0197] In this embodiment of the application, the feature data of the voice timbre, the first identity type, and the first behavior type can be input into a pre-trained sixth model to generate a first interactive voice that is used for interaction and has the voice timbre.

[0198] The sixth model can represent the correspondence between voice timbre, identity type, and behavior type. For example, the sixth model can be an artificial intelligence model trained using multiple sixth training samples, such as a large language model. The sixth training samples can include feature data of voice timbre, identity type, and behavior type.

[0199] As another example, a first interactive voice with the aforementioned timbre is generated through the process shown in Figure 11. Referring to Figure 11, a flowchart of another method for generating a first interactive voice with the aforementioned timbre is provided in an embodiment of this application. As shown in Figure 11, the process may include the following steps:

[0200] Step 1101: Obtain the seventh model corresponding to the above pronunciation timbre.

[0201] Step 1102: Input the first identity type and the first behavior type into the seventh model corresponding to the above voice timbre to obtain the first interactive voice that is used for interaction and has the above pronunciation timbre.

[0202] The following provides a unified explanation of steps 1101 and 1102:

[0203] In this embodiment, different voice timbres can correspond to different seventh models. Therefore, the first identity type and the first behavior type can be input into the seventh model corresponding to the voice timbre to generate a first interactive voice for interaction that possesses the stated voice timbre.

[0204] The seventh model can represent the correspondence between identity types and behavior types. For example, the seventh model can be an artificial intelligence model trained using training samples that include identity types and behavior types, such as a large language model.

[0205] As another example, text for interaction can be generated first based on a first identity type and a first behavior type. Then, a text-to-speech (TTS) model is used to generate a first interactive voice with the aforementioned voice timbre for interaction, based on the text and the aforementioned voice timbre.

[0206] For example, if the first identity type represents "suspicious person" and the first behavior type represents "loitering," then the first interactive voice could be a male voice (corresponding to the aforementioned voice tone) used to scare someone away. This helps ensure the safety of family members, especially when only children, the elderly, or women are present at home.

[0207] In some optional implementations of this embodiment, before determining the first voice interaction strategy matching the first visitor, the first voice interaction strategy can also be determined through the flowchart shown in FIG12. Referring to FIG12, it is a flowchart of a method for determining a first voice interaction strategy provided by an embodiment of this application. As shown in FIG12, the process may include the following steps:

[0208] Step 1201: Based on the first security video, determine the environmental information of the environment where the first visitor is located.

[0209] Step 1202: Based on at least one of the above environmental information, the above first identity type, and the above first behavior type, determine a first voice interaction strategy that matches the above first visitor.

[0210] The following provides a unified explanation of steps 1201 and 1202:

[0211] In this embodiment of the application, before determining the first voice interaction strategy that matches the first visitor, the environmental information of the environment in which the first visitor is located can also be determined based on the first security video.

[0212] Here, the environmental information of the environment in which the first visitor is located can be determined by identifying the first security video.

[0213] Based on this, the first voice interaction strategy matching the first visitor can be determined in the following way: based on at least one of the environmental information, the first identity type and the first behavior type, the first voice interaction strategy matching the first visitor can be determined.

[0214] As an example, at least one of the environmental information, the first identity type, and the first behavior type can be input into a pre-trained eighth model to determine a first voice interaction strategy that matches the first visitor.

[0215] The eighth model can represent the correspondence between at least one of environmental information, the first identity type, and the first behavior type, and the voice interaction strategy. As an example, the eighth model can be an artificial intelligence model trained using the eighth training sample, such as a large language model. The eighth training sample can include at least one of environmental information, the first identity type, and the first behavior type, as well as the voice interaction strategy.

[0216] Furthermore, at least two models among the first, second, third, fourth, fifth, sixth, seventh, and eighth models mentioned above can belong to the same model, or they can be independent models. In other words, at least two models among the first, second, third, fourth, fifth, sixth, seventh, and eighth models mentioned above can be a large model, or they can be several independent smaller models.

[0217] It is understandable that, among the above-mentioned optional implementation methods, different first voice interaction strategies can be determined for different environmental information, first identity type, or first behavior type. This can improve the matching degree between the first voice interaction strategy played by the security agent and the first security video, thereby improving security effectiveness.

[0218] In some optional implementations of this embodiment, the flowchart shown in Figure 13 can be used to determine the first identity type and first behavior type of the first visitor, and / or determine the first voice interaction strategy, and / or determine the first interactive voice. Referring to Figure 13, it is a flowchart of an embodiment of a voice interaction method in a security scenario provided by this application. As shown in Figure 13, the process may include the following steps:

[0219] Step 1301: Input the first security video into the multimodal big model to determine the first identity type and first behavior type of the first visitor.

[0220] Step 1302, and / or, using a multimodal large model, determine a first voice interaction strategy that matches the first visitor.

[0221] Step 1303, and / or, input the above-mentioned first voice interaction strategy, the above-mentioned first identity type and the first behavior type into the multimodal large model to generate the first interactive voice for interaction.

[0222] The following provides a unified explanation of steps 1301 to 1303:

[0223] In this embodiment of the application, the first identity type and first behavior type of the first visitor can be determined based on the first security video in the following way: the first security video is input into a multimodal big data model to determine the first identity type and first behavior type of the first visitor.

[0224] Here, the multimodal large model can be used to analyze video streams in real time, namely the first security video mentioned above.

[0225] Furthermore, the aforementioned multimodal large model, along with at least one of the aforementioned first, second, third, fourth, fifth, sixth, seventh, and eighth models, can be either a large model or an independent small model.

[0226] It is understandable that, among the above-mentioned optional implementation methods, a multimodal large model can be used to determine the first identity type and the first behavior type of the first visitor. This can improve the accuracy of determining the first identity type and the first behavior type. Furthermore, compared to other models, training a multimodal large model can improve the model's training efficiency, thereby improving the efficiency of determining the first identity type and the first behavior type.

[0227] In some optional implementations of this embodiment, the first voice interaction strategy matching the first visitor can be determined in the following way: by using a multimodal large model, the first voice interaction strategy matching the first visitor can be determined.

[0228] Here, the multimodal large model can determine the first voice interaction strategy that matches the first visitor.

[0229] It is understandable that, among the above-mentioned optional implementation methods, a multimodal large model can be used to determine the first voice interaction strategy that matches the first visitor. This can improve the matching degree between the determined first voice interaction strategy and the first visitor, and compared to other models, training a multimodal large model can improve the model's training efficiency, thereby improving the efficiency of determining the first voice interaction strategy.

[0230] In some optional implementations of this embodiment, the following method can be used to generate the first interactive voice for interaction based on the first identity type and the first behavior type, according to the first voice interaction strategy: input the first voice interaction strategy, the first identity type and the first behavior type into the multimodal large model to generate the first interactive voice for interaction.

[0231] Here, the multimodal large model can represent voice interaction strategies, identity types, behavior types, and the correspondence between interactive voices.

[0232] It is understandable that, among the above-mentioned optional implementation methods, a multimodal large model can be used to generate the first interactive speech for interaction. This can improve the accuracy of determining the first interactive speech, and compared with other models, training a multimodal large model can improve the training efficiency of the model, thereby improving the efficiency of generating the first interactive speech.

[0233] In some optional implementations of this embodiment, after the control security agent plays the first interactive voice, the steps shown in FIG14 can also be executed. Referring to FIG14, this is a flowchart of an embodiment of a voice interaction method in a security scenario provided by this application. As shown in FIG14, the process may include the following steps:

[0234] Step 1401: Obtain the recorded video of the security intelligent agent playing the first interactive voice.

[0235] The recorded video may be a video recorded by a security smart agent playing the first interactive voice, obtained by a video capture device such as a camera.

[0236] Step 1402: Determine the feedback information of the preset object to the recorded video.

[0237] The preset object can be an object that is pre-bound to the aforementioned execution entity and / or security intelligent agent. For example, the preset object can interact with the aforementioned execution entity and / or security intelligent agent through an application (APP). Thus, the preset object can be bound to the aforementioned execution entity and / or security intelligent agent through the account it logs into when using the application. For example, the preset object could be a homeowner.

[0238] Feedback information can indicate the preset object's level of satisfaction with the recorded video. For example, if the preset object likes the recorded video, the feedback information indicates that the preset object is relatively satisfied with the recorded video; if the preset object dislikes the recorded video, the feedback information indicates that the preset object is relatively dissatisfied with the recorded video.

[0239] Step 1403: Obtain the second security video containing the image of the second visitor.

[0240] The second visitor can be any one or more visitors. Visitors can be delivery personnel, police officers, neighbors, or other individuals; they can also be mobile electronic devices such as drones or vehicles; or even pets. The second visitor can be the same as or different from the first visitor.

[0241] The second security video can be a video containing images of a second visitor. This second security video can be acquired via a camera. For example, the aforementioned camera can be installed at building entrances, corridors, elevator lobbies, stairwells, etc. Alternatively, the aforementioned camera can be installed on a door, such as a door lock or doorbell.

[0242] In some cases, the first security video and the second security video can be captured by the same camera.

[0243] Step 1404: Based on the second security video, determine the second identity type and second behavior type of the second visitor.

[0244] The second identity type can represent the identity of the second visitor. For example, if the second visitor is a neighbor, the second identity type can represent the neighbor; if the second visitor is a deliveryman, the second identity type can represent the deliveryman; if the second visitor is the homeowner, the second identity type can represent the homeowner; and if the identity of the second visitor is not identified, the second identity type can represent a stranger.

[0245] The second behavior type can represent the behavior of the second visitor. For example, if the second visitor's behavior is delivering a package, the second behavior type can represent delivering a package; if the second visitor's behavior is loitering, the second behavior type can represent loitering.

[0246] Here, multiple methods can be used to determine the second identity type of the second visitor based on the second security video.

[0247] As an example, a second security video can be input into a pre-trained first model to determine the second identity type of the second visitor. The meaning of the first model is explained above and will not be repeated here.

[0248] As another example, images of objects with different identity types can be acquired first, thus obtaining a first image set. This first image set may include multiple acquired images, and each image in the first image set may correspond to one identity type. Next, feature data from the second security video can be extracted, and the similarity between this feature data and the feature data of each image in the first image set can be calculated. Then, the identity type corresponding to the highest similarity value among the calculated similarities is determined as the second identity type of the second visiting object.

[0249] The feature data in the above examples may include, but is not limited to, facial features, iris features, fingerprint features, gait features, etc.

[0250] In addition, the similarity in the above examples can be obtained by calculating Euclidean distance, cosine similarity, and Hamming distance.

[0251] In addition, various methods can be used to determine the second behavior type of the second visitor based on the second security video.

[0252] As an example, a second security video can be input into a pre-trained second model to determine the second behavior type of the second visitor. The meaning of the representation in the second model is described above and will not be repeated here.

[0253] As another example, images of objects with different behavior types can be acquired first to obtain a second image set. This second image set may include multiple acquired images, and each image in the second image set may correspond to a behavior type. Next, feature data from the second security video can be extracted, and the similarity between this feature data and the feature data of each image in the second image set can be calculated. Then, the behavior type corresponding to the highest similarity value among the calculated similarities is determined as the second behavior type of the second visiting object.

[0254] The feature data in the above examples may include, but is not limited to, joint features, limb features, etc.

[0255] In addition, the similarity in the above examples can be obtained by calculating Euclidean distance, cosine similarity, and Hamming distance.

[0256] In some cases, steps one and two can be executed in parallel, or step one can be executed first, and step two can be executed only if the second identity type determined in step one is a preset identity type. The preset identity type can represent a stranger.

[0257] The first step is to determine the second identity type of the second visitor based on the second security video. The second step is to determine the second behavior type of the second visitor based on the second security video.

[0258] Step 1405: Based on the feedback information, determine a second voice interaction strategy that matches the second visitor.

[0259] The second voice interaction strategy can be a voice interaction strategy that matches the second visitor.

[0260] The voice interaction strategy (including the second voice interaction strategy) may include at least one of the following: the pronunciation style of the speech object (e.g., polite, humorous, stern, friendly) and the timbre of the speech object (e.g., the timbre of an adult male, the timbre of a child, etc.).

[0261] Here, various methods can be used to determine a second voice interaction strategy that matches the second visitor based on the feedback information.

[0262] In some optional implementations of this embodiment, before determining the second voice interaction strategy matching the second visitor based on the feedback information, the following steps may also be performed:

[0263] The first step is to identify the interviewees of the second visitor.

[0264] In this context, the recipients of the second visitor's interview can be any of the individuals or entities within the location visited by the second visitor. Alternatively, the recipients of the second visitor's interview can also be identified through the second visitor's voice. For example, if the second visitor utters the voice message "Is XX home?", then it can be determined that the recipient of the second visitor's interview is XX.

[0265] The second step is to determine the object information of the interviewee.

[0266] The object information includes at least one of the following: the interviewee's voice timbre, the interviewee's preference information, the interviewee's age, the interviewee's gender, and the number of interviewees.

[0267] The preference information of the respondents may include at least one of the following: the voice interaction strategy set by the respondents for the respondents’ identity type and behavior type, the voice timbre set by the respondents, etc.

[0268] Based on this, a second voice interaction strategy matching the second visitor can be determined using the following method, based on the feedback information: a second voice interaction strategy matching the second visitor can be determined based on the object information and the feedback information.

[0269] For example, the object information and the feedback information can be input into a pre-trained multimodal large model to determine a second voice interaction strategy that matches the second visiting object.

[0270] The aforementioned multimodal large model can represent the correspondence between object information, feedback information, and voice interaction strategies. This multimodal large model can be trained using training samples that include object information, feedback information, and voice interaction strategies.

[0271] In some optional implementations of this embodiment, before determining the second voice interaction strategy matching the second visitor based on the feedback information, the environmental information of the environment in which the second visitor is located can also be determined based on the second security video.

[0272] Based on this, a second voice interaction strategy matching the second visitor can be determined using the feedback information, including:

[0273] Based on at least one of the environmental information, the second identity type, and the second behavior type, as well as the feedback information, a second voice interaction strategy matching the second visitor is determined.

[0274] As an example, at least one of the environmental information, the second identity type, and the second behavior type, as well as the feedback information, can be input into a pre-trained multimodal large model to determine a second voice interaction strategy that matches the second visitor.

[0275] Step 1406: In accordance with the second voice interaction strategy, based on the second identity type and the second behavior type, generate a second interactive voice for interacting with the second visitor.

[0276] The second interactive voice can conform to the aforementioned second voice interaction strategy. For example, if the second voice interaction strategy indicates that the pronunciation style is polite, the second interactive voice can be polite, such as "Hello, please leave the package at the door." If the second voice interaction strategy indicates that the pronunciation style is an adult male's tone, the second interactive voice can have that adult male's tone, such as saying "Leave as soon as possible" in a stern adult male's tone.

[0277] Here, various methods can be used to generate the second interactive voice for interaction based on the second identity type and the second behavior type, according to the second voice interaction strategy. For details, please refer to the description in the context of generating the first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy; it will not be repeated here.

[0278] Step 1407: Control the security intelligent agent to play the second interactive voice.

[0279] Here, after generating the second interactive voice, the security agent can be further controlled to play the second interactive voice.

[0280] It is understandable that, among the above-mentioned optional implementation methods, the voice interaction strategy can be dynamically adjusted to match subsequent visitors based on the feedback information from the preset visitors to the recorded video. This can improve the satisfaction of the preset visitors with the voice interaction strategy, thereby making the security strategy more aligned with the preset visitors' security intentions.

[0281] It should be noted that, where there is no conflict, the technical features described in different alternative implementations can be included in the same embodiment. For the sake of brevity, they will not be elaborated here.

[0282] The voice interaction method in a security scenario provided in this application embodiment can acquire a first security video containing an image of a first visitor. Then, based on the first security video, a first identity type and a first behavior type of the first visitor are determined. Next, a first voice interaction strategy matching the first visitor is determined. Subsequently, according to the first voice interaction strategy, a first interactive voice for interaction is generated based on the first identity type and the first behavior type. Finally, the security intelligent agent is controlled to play the first interactive voice. Therefore, according to the voice interaction strategy matching the visitor, and based on the visitor's identity type and behavior type, interactive voice for interaction can be generated, and the security intelligent agent can be controlled to play the interactive voice, thus improving the security of visitors in a security scenario.

[0283] Figure 15 is a flowchart illustrating another voice interaction method in a security scenario provided by an embodiment of this application. As shown in Figure 15, the method specifically includes:

[0284] Step 1501: Obtain the first security video containing the image of the first visitor.

[0285] In this embodiment, step 1501 is basically the same as step 101 in the embodiment corresponding to Figure 1, and will not be described again here.

[0286] Step 1502: Based on the first security video, determine the first identity type and first behavior type of the first visitor.

[0287] In this embodiment, step 1502 is basically the same as step 102 in the embodiment corresponding to Figure 1, and will not be described again here.

[0288] Step 1503: Determine the interviewee of the first visitor.

[0289] In this embodiment, the recipient of the first visitor's visit can be any of the objects included in the location visited by the first visitor. Alternatively, the recipient of the first visitor's visit can also be determined by the first visitor's voice. For example, if the first visitor says "Is YY home?", then it can be determined that the recipient of the first visitor's visit is YY.

[0290] Step 1504: Determine the object information of the interviewees, wherein the object information includes at least one of the following: the interviewee's voice timbre, the interviewee's preference information, the interviewee's age, the interviewee's gender, and the number of interviewees.

[0291] In this embodiment, the preference information of the interviewee may include at least one of the following: the voice interaction strategy set by the interviewee for the interviewee's identity type and behavior type, the timbre set by the interviewee, etc.

[0292] Step 1505: Based on the object information, determine a first voice interaction strategy that matches the first visiting object.

[0293] In this embodiment, a first voice interaction strategy matching the first visiting object can be determined in a variety of ways based on the object information.

[0294] For example, object information can be input into a pre-trained multimodal large model to determine a first voice interaction strategy that matches the first visiting object.

[0295] The aforementioned multimodal large model can represent the correspondence between object information and voice interaction strategies. This multimodal large model can be trained using training samples that include both object information and voice interaction strategies.

[0296] Step 1506: In accordance with the first voice interaction strategy, generate the first interactive voice for interaction based on the first identity type and the first behavior type.

[0297] In this embodiment, step 1506 is basically the same as step 104 in the embodiment corresponding to Figure 1, and will not be described again here.

[0298] Step 1507: Based on the object information, determine a first voice interaction strategy that matches the first visiting object.

[0299] In this embodiment, step 1507 is basically the same as step 105 in the embodiment corresponding to Figure 1, and will not be described again here.

[0300] In some optional implementations of this embodiment, when the object information includes the preference information of the visited object, the process shown in FIG16 can be adopted to determine a first voice interaction strategy matching the first visited object based on the object information. Referring to FIG16, which is a flowchart of a method for generating a first voice interaction strategy provided by an embodiment of this application, as shown in FIG16, the process may include the following steps:

[0301] Step 1601: Determine whether the preference information includes the voice interaction strategy triggered by the first visitor.

[0302] Step 1602: If the preference information includes the voice interaction strategy triggered by the first visitor, the voice interaction strategy triggered by the first visitor is determined as the first voice interaction strategy that matches the first visitor.

[0303] Step 1603: If the preference information does not include the voice interaction strategy triggered by the first visitor, generate a first voice interaction strategy that matches the first visitor based on the preference information.

[0304] As an example, if the preference information includes a predefined voice interaction strategy, it will be executed directly. For instance, if the predefined target is set to prompt for a message when a neighbor visits, a message command can be output, triggering the message function via hardware. If the preference information does not include a predefined voice interaction strategy, a voice interaction strategy that conforms to the predefined target's past practices can be autonomously generalized based on previous voice interaction strategies. For example, if the predefined target has not specified how to respond to a delivery person's visit, it may proactively decide based on previous voice interaction strategies and output a message thanking the delivery person.

[0305] It is understood that, among the above-mentioned optional implementation methods, generating a first voice interaction strategy that matches the first visitor based on the preference information can improve the applicability of the solution.

[0306] It should be noted that, in addition to the contents described above, this embodiment may also include the corresponding technical features described in the embodiment corresponding to Figure 1, thereby achieving the technical effect of the voice interaction method in the security scenario shown in Figure 1. For details, please refer to the relevant description in Figure 1. For the sake of brevity, it will not be elaborated here.

[0307] The voice interaction method in the security scenario provided in this application embodiment can determine a first voice interaction strategy that matches the first visitor based on at least one of the visitor's vocal timbre, the visitor's preference information, the visitor's age, the visitor's gender, and the number of visitors. This improves the matching degree between the first voice interaction strategy and the first visitor.

[0308] The embodiments of this application are described below by way of example. However, it should be noted that the embodiments of this application may have the features described below, but the following description does not constitute a limitation on the protection scope of the embodiments of this application.

[0309] Before introducing this solution, the technical terms involved in this solution will be explained as follows:

[0310] Large language models are natural language processing models based on deep learning techniques, capable of understanding, generating, and processing human language. They are trained using massive amounts of text data and learn the syntax, semantics, and contextual relationships of a language through complex neural network architectures, such as the Transformer model.

[0311] A multimodal model is a model that processes and understands multiple data types (such as text, images, audio, and video). In this scheme, it is used for real-time analysis of video streams (e.g., first security video, second security video) and person states (i.e., the aforementioned first identity type, first behavior type, second identity type, and second behavior type).

[0312] Text-to-speech (TS) is a technology that converts text content into speech output, enabling devices to "speak".

[0313] Behavior recognition is a technique that uses multimodal large models to determine the type of behavior of a person by analyzing their actions and environment.

[0314] A feature vector is a vector used in machine learning and data analysis to represent the characteristics of data. It transforms raw data into a set of numerical features so that algorithms can process and analyze it. These features are quantitative descriptions of the original data and can be any type of data, such as images, text, audio, etc.

[0315] The process of constructing feature vectors typically includes data preprocessing, feature extraction, and feature selection. Data preprocessing may involve noise removal, standardization, and normalization. Feature extraction is the process of converting data into feature vectors, and the specific method depends on the type of data. In image processing, pixel values, edge detection results, or the output of intermediate layers in a deep learning model can all serve as feature vectors. In machine learning models, feature vectors represent the input data, and the model makes predictions or classifications by learning the mapping relationship between feature vectors and the target output.

[0316] In related technologies, home doorbells typically only have functions such as monitoring the doorway area, detecting packages, and detecting and recognizing people. As for dialogue, they mainly rely on an app or indoor voice devices to allow the homeowner to have a real-time conversation with visitors at the door. Overall, the functions are relatively limited, and the conversation with visitors cannot be conducted without the homeowner's real-time participation, resulting in a poor overall user experience.

[0317] Existing smart doorbells have the following technical problems:

[0318] Limited functionality and low level of intelligence: It cannot deeply understand the identity and behavior of visitors and lacks personalized response capabilities.

[0319] It relies on real-time user participation: the owner needs to communicate with visitors at the door in real time, and it cannot respond autonomously when the owner is not present or cannot respond in a timely manner.

[0320] Limited security protection capabilities: There is a lack of proactive warnings or dissuasion measures for suspicious individuals, making it impossible to effectively protect user safety.

[0321] Insufficient support for special groups: There is a lack of thoughtful functional design for the elderly, children, and women living alone, which fails to meet their special needs.

[0322] In view of this, this solution proposes a method combining a multimodal large model to transform the smart doorbell into a security intelligent agent. By combining the user's (i.e., the aforementioned preset object) preferences (i.e., the aforementioned preference information) settings with the large model's learning and summarization of user habits, the doorbell can receive visitors (i.e., the aforementioned visitors) according to the owner's (i.e., the aforementioned preset object) settings and preferences. At the same time, when the large model identifies suspicious persons such as thieves, drunkards, or loitering at the door (i.e., the aforementioned first identity type and second identity type), it can simulate different timbres and words to drive them away and deter them, further protecting the user's safety. In particular, it can greatly improve the safety of some vulnerable groups (such as the elderly, children, and women) who are home alone.

[0323] Overall, existing smart doorbells have many limitations in terms of functionality and user experience, failing to meet users' higher demands for home security and convenient interaction. Therefore, this solution introduces a multimodal large model and a large language model to enhance the intelligence level of the smart doorbell, enabling it to have proactive recognition, personalized response, and adaptive learning capabilities, thus achieving a true security intelligence agent and providing a safer, more considerate, and convenient user experience.

[0324] This plan involves the following entities:

[0325] 1. Smart doorbell: Equipped with a camera, microphone, speaker, and necessary computing power, it can run multimodal models and large language models.

[0326] 2. Mobile devices (such as smartphones): These devices have a dedicated app installed, allowing users to set up devices, receive notifications, and view the situation in front of the doorbell.

[0327] 3. Central Control System: Integrates data from all smart devices, performs unified analysis, decision-making, and control, and interacts with users through a user app.

[0328] 4. User APP: Used to receive information pushed by the system, view real-time monitoring videos and historical records, interact with the system, and issue alarms when necessary.

[0329] The application scenarios of this solution include, but are not limited to:

[0330] 1. Typical family environment: Suitable for all family homes, especially those with elderly people, children, or members living alone who require a higher level of security.

[0331] 2. When visitors arrive:

[0332] Neighbor visit: The smart doorbell recognizes a neighbor and responds politely or humorously, according to the user's settings.

[0333] Delivery by courier: Identify the courier, instruct them to place the package in the designated area, and express your gratitude to them.

[0334] 3. Family members return home:

[0335] When a child comes home: The smart doorbell recognizes the child, proactively sends a notification to the parents, and greets them with a warm voice.

[0336] 4. Suspicious individuals appear:

[0337] Stranger loitering: When a stranger is detected loitering at the door, the smart doorbell will issue a warning in a stern tone, urging the stranger to leave.

[0338] Thief or drunkard: When a potential threat is identified, the security mechanism is automatically triggered, an alarm is sent to the user, and video recording or other security measures may be initiated.

[0339] 5. The user is not at home or cannot respond in real time:

[0340] Automatic visitor reception: When users are unable to participate in the conversation in real time, the smart doorbell interacts with visitors on their behalf according to preset preferences, ensuring that important information is not missed.

[0341] 6. Special needs scenarios:

[0342] Elderly and vulnerable groups living alone: ​​Provides extra safety and thoughtful interaction for the elderly, children, or women living alone at home.

[0343] This solution applies the software / algorithm / method to smart doorbells, enabling the device to intelligently identify information such as the scene at the door, the identity of the person (i.e., the first identity type and the second identity type mentioned above), and the person's behavior (i.e., the first behavior type and the second behavior type mentioned above). Based on the user's preferences, it provides personalized interaction and security measures, thereby improving home security and the user experience.

[0344] Specifically, referring to Figures 17A and 17B, this scheme includes the following steps:

[0345] 1. User settings: Input of character information.

[0346] During the setup process, users can enable the smart doorbell to recognize newly added family members. For example, a person can stand 1 meter away from the doorbell and engage in a pre-set dialogue to complete the entry of the person's information. The background model can learn the person's appearance features (such as facial features) and corresponding voice timbre. Then, the app will present the corresponding virtual person. After the user confirms the relationship between the members, when the female homeowner clicks on the virtual person on her phone, the virtual person will respond with the corresponding member's voice timbre. For example, clicking on the child will prompt the app to say "Hi, Mommy" in the child's voice.

[0347] Clicking on a character card in the app will trigger a corresponding voice output based on the user's identity. For example, if the user is the female homeowner, clicking on her child's character card will produce a corresponding output, such as "Hi, Mommy".

[0348] 2. User Settings: Preferences.

[0349] The app interface presents users with options for different scenarios, such as: a neighbor visiting, a courier delivering a package, a stranger loitering, a child returning home, etc. For different scenarios, some options can be preset for users. For example, when a neighbor visits, the intelligent agent can be set to respond in the female homeowner's voice (i.e., the first and second voice interaction strategies mentioned above). When a child's friend visits, the agent can be set to respond in the child's voice, and the pronunciation style of the response can be set, such as: polite version, humorous version, etc.

[0350] 3. Scene recognition.

[0351] Objective: To analyze entrance video in real time to identify visitors and the context of their visit.

[0352] Implementation steps:

[0353] When the doorbell is triggered, the camera begins recording a video stream. A multimodal big data model analyzes the video footage in real time, extracting human features. It identifies a person's family attributes (e.g., family members, face matching), occupational characteristics (e.g., deliveryman uniform), or behavioral characteristics (e.g., loitering, face concealment), and then passes the identification results to the behavior determination module.

[0354] 4. Behavior determination.

[0355] When the doorbell receives a video stream, it extracts frames from the video stream and sends them to a multimodal large model. The vision module of the large model extracts vision embeddings (visual feature vectors) from the image frames. After the video stream continues to be input, multimodal inference begins when T frames (T is a preset positive integer) are satisfied. The specific internal logic is as follows: T vision embeddings are concatenated and input into the multimodal large model, and the behavior category (i.e., the behavior type mentioned above) and behavior description are output. The whole process is carried out in real time, that is, after T frames are satisfied, the large model inference is performed every second, thereby ensuring the timeliness of behavior recognition.

[0356] 5. Intelligent dialogue: thoughtful care and safety protection.

[0357] 5.1 Custom Voice Generation: The doorbell has a built-in TTS model that supports inputting a user's voice and mapping the voice of the subsequent text-to-speech conversion to the user's voice to achieve the effect of a real human voice.

[0358] 5.2. Intent judgment.

[0359] 5.2.1. The multimodal large model includes a vision module and a large language model module. The large language module can respond and reply according to different language inputs.

[0360] 5.2.2. User Preference Vector Library: All user preference settings are input into the large model through a special text format. The large model extracts them into a vector library and stores it as a knowledge base of user preferences.

[0361] 5.2.3. After the behavior is determined, the corresponding behavior category and environmental information are transmitted back to the large model. The large model then combines this with a user preference vector library to output corresponding responses. For example, if the user library contains user-specified decisions, they are executed directly. For instance, if the user has set a message prompt when a neighbor visits, the large model will output a message instruction, triggering the message function via hardware. If the user library does not contain user-specified rules, the large model will combine past rules to autonomously generalize behavioral decisions that conform to the user's past habits. For example, if the user has not set how to respond to a courier's visit, the system may proactively decide based on past preferences and output a thank-you message to the courier.

[0362] 5.2.4. Intent determination is one of the core functions of the system's overall model. It combines user preferences to respond to different behaviors, output different commands, and schedule the functions of the hardware on the doorbell. As mentioned above, it can also be combined with TTS (Text-to-Speech) so that when encountering suspicious persons, the doorbell can mimic the homeowner's aggressive tone to drive them away.

[0363] 6. Adaptive user learning.

[0364] 6.1. Two-way fitting: In the early stage, based on the user's set preferences, the smart doorbell will conduct corresponding scene dialogues and greet visitors. At the same time, it will record each video, and the likes and dislikes will be used as key practices and pushed to the user. The user can participate in rating and evaluating it. If the smart agent does not handle the situation well, it can also give corresponding reasonable evaluations and guidance. The smart agent will continuously learn the user's habits and style and correct the next processing method, thus realizing the process of two-way fitting between the user and the smart agent.

[0365] 6.2. Targeted Empowerment by the App: The user habits and styles learned by the large model will be presented in a structured form on the app, similar to the way the user set their own preferences. Users can see the doorbell's handling style in different scenarios in real time. Of course, users can also directly modify the corresponding style to achieve the desired effect.

[0366] It should be noted that, in addition to the contents described above, this embodiment may also include the technical features described in the above embodiments, thereby achieving the technical effect of the voice interaction method in the security scenario shown above. Please refer to the above description for details. For the sake of brevity, it will not be elaborated here.

[0367] The voice interaction method in the security scenario provided in this application combines a multimodal large model with a smart doorbell to create an intelligent agent that understands the user. This agent can recognize information about people and their behaviors in videos, and combined with voice capabilities, can greet visitors, care for family members (elderly and children, etc.), protect property, ensure safety, and more effectively deter dangerous individuals. Through user preference settings and a personalized response mechanism, the app allows users to set different response methods for different scenarios, including voice selection (e.g., lady of the house, child) and dialogue style (polite version, humorous version). The smart doorbell adaptively interacts with visitors based on these preferences. This achieves highly personalized visitor reception, meets diverse user needs, and enhances the user experience. By recording each visitor interaction and user feedback, the smart doorbell's agent can continuously learn and optimize its behavior and response methods, achieving bidirectional fitting with user habits and styles. This enables adaptive learning and bidirectional fitting of the agent. The introduction of an adaptive learning mechanism allows the device to continuously improve and better meet user needs. By introducing multimodal and large language models, the smart doorbell is upgraded into a security agent with advanced intelligent interaction and proactive security protection capabilities. Furthermore, this solution enhances multimodal perception capabilities, enabling accurate identification of visitor identity and behavior, and providing more intelligent scene understanding. In terms of personalized response, it can automatically adapt voice tone and dialogue style based on user preferences for personalized interaction with visitors. Regarding proactive security protection, it can provide real-time alerts and dissuade suspicious individuals, proactively improving home security. In terms of adaptive learning capabilities, it can continuously optimize the agent's behavior based on user feedback, better aligning with user habits and needs. It can also provide support for special groups, such as offering considerate protection and interaction methods for the elderly, children, and users home alone. Regarding the two-way fitting interactive experience, users can view and adjust the agent's behavior in real time through the app, enhancing their sense of control over the device. It provides a convenient and intuitive settings and management interface, enhancing the user experience and seamlessly integrating app functions. These advantages significantly improve the intelligence level and home security protection capabilities of the smart doorbell, meeting users' higher requirements for personalization, security, and convenience.

[0368] Figure 18 is a structural schematic diagram of a voice interaction device in a security scenario provided by an embodiment of this application. Specifically, it includes:

[0369] The first acquisition unit 1801 is used to acquire a first security video containing an image of a first visitor.

[0370] The first determining unit 1802 is used to determine the first identity type and the first behavior type of the first visitor based on the first security video.

[0371] The second determining unit 1803 is used to determine a first voice interaction strategy that matches the first visitor.

[0372] The first generation unit 1804 is configured to generate a first interactive voice for interaction based on the first identity type and the first behavior type, according to the first voice interaction strategy.

[0373] The first control unit 1805 is used to control the security intelligent agent to play the first interactive voice.

[0374] In one possible implementation, determining the first identity type of the first visitor based on the first security video includes:

[0375] The first security video is input into a pre-trained first model to obtain the first identity type of the first visitor.

[0376] In one possible implementation, determining the first identity type of the first visitor based on the first security video includes:

[0377] Obtain images of objects of different identity types to obtain the first image set;

[0378] Extract the feature data of the first security video;

[0379] Calculate the similarity between the feature data and the feature data of each image in the first image set;

[0380] The identity type corresponding to the highest similarity value among the obtained similarities is determined as the first identity type of the first visitor.

[0381] In one possible implementation, based on the first security video, determining the first behavior type of the first visitor includes:

[0382] The first security video is input into a pre-trained second model to obtain the first behavior type of the first visitor.

[0383] In one possible implementation, based on the first security video, determining the first behavior type of the first visitor includes:

[0384] Obtain images of objects with different behavior types to obtain a second image set;

[0385] Extract the feature data of the first security video;

[0386] Calculate the similarity between the feature data and the feature data of each image in the second image set;

[0387] The behavior type corresponding to the highest similarity value among the obtained similarities is determined as the first behavior type of the first visitor.

[0388] In one possible implementation, the first voice interaction strategy includes the pronunciation style of a single speaker during voice interaction; and

[0389] The step of generating a first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy includes:

[0390] Based on the first identity type and the first behavior type, a first interactive voice with the pronunciation style is generated for interaction.

[0391] In one possible implementation, generating a first interactive voice with the pronunciation style based on the first identity type and the first behavior type includes:

[0392] Obtain the feature data of the pronunciation style;

[0393] The feature data of the pronunciation style, the first identity type, and the first behavior type are input into a pre-trained fourth model to obtain a first interactive voice with the pronunciation style for interaction.

[0394] In one possible implementation, generating a first interactive voice with the pronunciation style based on the first identity type and the first behavior type includes:

[0395] Obtain the fifth model corresponding to the pronunciation style;

[0396] The first identity type and the first behavior type are input into the fifth model corresponding to the pronunciation style to obtain a first interactive voice that is used for interaction and has the pronunciation style.

[0397] In one possible implementation, the first voice interaction strategy includes the voice timbre for performing the voice interaction; and

[0398] The step of generating a first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy includes:

[0399] Based on the first identity type and the first behavior type, a first interactive voice with the aforementioned voice timbre is generated for interaction.

[0400] In one possible implementation, generating a first interactive voice for interaction and having the pronunciation timbre based on the first identity type and the first behavior type includes:

[0401] Obtain the characteristic data of the pronunciation timbre;

[0402] The feature data of the pronunciation timbre, the first identity type, and the first behavior type are input into the pre-trained sixth model to obtain the first interactive voice with the pronunciation timbre for interaction.

[0403] In one possible implementation, generating a first interactive voice for interaction and having the pronunciation timbre based on the first identity type and the first behavior type includes:

[0404] Obtain the seventh model corresponding to the pronunciation timbre;

[0405] The first identity type and the first behavior type are input into the seventh model corresponding to the voice timbre to obtain a first interactive voice that is used for interaction and has the pronunciation timbre.

[0406] In one possible implementation, before determining the first voice interaction strategy matching the first visitor, the device further includes:

[0407] The third determining unit (not shown in the figure) is used to determine the interviewee of the first visiting object;

[0408] The fourth determining unit (not shown in the figure) is used to determine the object information of the interviewees, wherein the object information includes at least one of the following: the interviewee's voice timbre, the interviewee's preference information, the interviewee's age, the interviewee's gender, and the number of interviewees; and

[0409] The first voice interaction strategy for determining the match with the first visitor includes:

[0410] Based on the object information, a first voice interaction strategy matching the first visiting object is determined.

[0411] In one possible implementation, when the object information includes the preference information of the interviewee, determining a first voice interaction strategy matching the first interviewee based on the object information includes:

[0412] Determine whether the preference information includes a voice interaction strategy triggered by the first visitor;

[0413] If the preference information includes a voice interaction strategy triggered by the first visitor, the voice interaction strategy triggered by the first visitor is determined as a first voice interaction strategy that matches the first visitor.

[0414] If the preference information does not include the voice interaction strategy triggered by the first visitor, a first voice interaction strategy matching the first visitor is generated based on the preference information.

[0415] In one possible implementation, before determining the first voice interaction strategy matching the first visitor, the device further includes:

[0416] The fifth determining unit (not shown in the figure) is used to determine the environmental information of the environment where the first visitor is located based on the first security video; and

[0417] The first voice interaction strategy for determining the match with the first visitor includes:

[0418] Based on at least one of the environmental information, the first identity type, and the first behavior type, a first voice interaction strategy matching the first visitor is determined.

[0419] In one possible implementation,

[0420] The step of determining the first identity type and first behavior type of the first visitor based on the first security video includes:

[0421] The first security video is input into a multimodal big data model to determine the first identity type and first behavior type of the first visitor; and / or

[0422] The first voice interaction strategy for determining the match with the first visitor includes:

[0423] Using a multimodal large model, determine a first voice interaction strategy that matches the first visitor; and / or

[0424] The step of generating a first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy includes:

[0425] The first voice interaction strategy, the first identity type, and the first behavior type are input into the multimodal large model to generate the first interactive voice for interaction.

[0426] In one possible implementation, after the security agent plays the first interactive voice, the device further includes:

[0427] The second acquisition unit (not shown in the figure) is used to acquire the recorded video of the security intelligent agent playing the first interactive voice.

[0428] The sixth determining unit (not shown in the figure) is used to determine the feedback information of the preset object to the recorded video;

[0429] The third acquisition unit (not shown in the figure) is used to acquire a second security video containing an image of the second visitor.

[0430] The seventh determining unit (not shown in the figure) is used to determine the second identity type and the second behavior type of the second visitor based on the second security video.

[0431] The eighth determining unit (not shown in the figure) is used to determine a second voice interaction strategy that matches the second visitor based on the feedback information;

[0432] The second generation unit (not shown in the figure) is used to generate a second interactive voice for interacting with the second visitor based on the second identity type and the second behavior type, according to the second voice interaction strategy.

[0433] The second control unit (not shown in the figure) is used to control the security smart agent to play the second interactive voice.

[0434] In one possible implementation, the security agent is a smart doorbell.

[0435] The voice interaction device in the security scenario provided in this embodiment can be the voice interaction device in the security scenario shown in Figure 4. It can execute all the steps of the voice interaction methods in the various security scenarios described above, thereby achieving the technical effects of the voice interaction methods in the various security scenarios described above. For details, please refer to the relevant descriptions above. For the sake of brevity, it will not be elaborated here.

[0436] Figure 19 is a schematic diagram of an electronic device provided in an embodiment of this application. The electronic device 1900 shown in Figure 19 includes: at least one processor 1901, a memory 1902, at least one network interface 1904, and other user interfaces 1903. The various components in the electronic device 1900 are coupled together through a bus system 1905. It is understood that the bus system 1905 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 1905 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 1905 in Figure 19.

[0437] The user interface 1903 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).

[0438] It is understood that the memory 1902 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 1902 described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0439] In some implementations, memory 1902 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 19021 and application program 19022.

[0440] The operating system 19021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 19022 includes various applications, such as a media player and a browser, used to implement various application functions. Programs implementing the methods of the embodiments of this application can be included in application program 19022.

[0441] In this embodiment, by calling the program or instructions stored in memory 1902, specifically the program or instructions stored in application program 19022, processor 1901 executes the method steps provided in each method embodiment, including, for example:

[0442] Acquire the first security video containing images of the first visitor;

[0443] Based on the first security video, determine the first identity type and first behavior type of the first visitor;

[0444] Determine a first voice interaction strategy that matches the first visitor;

[0445] According to the first voice interaction strategy, a first interactive voice for interaction is generated based on the first identity type and the first behavior type;

[0446] Control the security intelligent agent to play the first interactive voice.

[0447] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1901. The processor 1901 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 1901 or by instructions in the form of software. The processor 1901 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 1902. Processor 1901 reads the information in memory 1902 and, in conjunction with its hardware, completes the steps of the above method.

[0448] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described above, or combinations thereof.

[0449] For software implementation, the techniques described herein can be implemented by units that perform the functions described above. The software code can be stored in memory and executed by a processor. The memory can be implemented within the processor or external to the processor.

[0450] The electronic device provided in this embodiment can be the electronic device shown in Figure 19, which can execute all the steps of the voice interaction methods in the various security scenarios described above, thereby achieving the technical effects of the voice interaction methods in the various security scenarios described above. For details, please refer to the relevant descriptions above. For the sake of brevity, it will not be elaborated here.

[0451] This application also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; it may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; and it may also include combinations of the above types of memory.

[0452] When one or more programs in the storage medium can be executed by one or more processors to implement the aforementioned voice interaction method in security scenarios executed on the electronic device side.

[0453] The aforementioned processor is used to execute a voice interaction program for a security scenario stored in memory, in order to implement the following steps of the voice interaction method for a security scenario executed on the electronic device side:

[0454] Acquire the first security video containing images of the first visitor;

[0455] Based on the first security video, determine the first identity type and first behavior type of the first visitor;

[0456] Determine a first voice interaction strategy that matches the first visitor;

[0457] According to the first voice interaction strategy, a first interactive voice for interaction is generated based on the first identity type and the first behavior type;

[0458] Control the security intelligent agent to play the first interactive voice.

[0459] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0460] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0461] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0462] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A voice interaction method in a security scenario, characterized in that, The method includes: Acquire the first security video containing images of the first visitor; Based on the first security video, determine the first identity type and first behavior type of the first visitor; Determine a first voice interaction strategy that matches the first visitor; According to the first voice interaction strategy, a first interactive voice for interaction is generated based on the first identity type and the first behavior type; Control the security intelligent agent to play the first interactive voice.

2. The method according to claim 1, characterized in that, Based on the first security video, the first identity type of the first visitor is determined, including: The first security video is input into a pre-trained first model to obtain the first identity type of the first visitor.

3. The method according to claim 1, characterized in that, Based on the first security video, the first identity type of the first visitor is determined, including: Obtain images of objects of different identity types to obtain the first image set; Extract the feature data of the first security video; Calculate the similarity between the feature data and the feature data of each image in the first image set; The identity type corresponding to the highest similarity value among the obtained similarities is determined as the first identity type of the first visitor.

4. The method according to claim 1, characterized in that, Based on the first security video, the first behavior type of the first visitor is determined, including: The first security video is input into a pre-trained second model to obtain the first behavior type of the first visitor.

5. The method according to claim 1, characterized in that, Based on the first security video, the first behavior type of the first visitor is determined, including: Obtain images of objects with different behavior types to obtain a second image set; Extract the feature data of the first security video; Calculate the similarity between the feature data and the feature data of each image in the second image set; The behavior type corresponding to the highest similarity value among the obtained similarities is determined as the first behavior type of the first visitor.

6. The method according to claim 1, characterized in that, The first voice interaction strategy includes the pronunciation style of the pronunciation object when performing voice interaction; as well as The step of generating a first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy includes: Based on the first identity type and the first behavior type, a first interactive voice with the pronunciation style is generated for interaction.

7. The method according to claim 6, characterized in that, The step of generating a first interactive voice with the pronunciation style for interaction based on the first identity type and the first behavior type includes: Obtain the feature data of the pronunciation style; The feature data of the pronunciation style, the first identity type, and the first behavior type are input into a pre-trained fourth model to obtain a first interactive voice with the pronunciation style for interaction.

8. The method according to claim 6, characterized in that, The step of generating a first interactive voice with the pronunciation style for interaction based on the first identity type and the first behavior type includes: Obtain the fifth model corresponding to the pronunciation style; The first identity type and the first behavior type are input into the fifth model corresponding to the pronunciation style to obtain a first interactive voice that is used for interaction and has the pronunciation style.

9. The method according to claim 1, characterized in that, The first voice interaction strategy includes the pronunciation timbre for performing voice interaction; as well as The step of generating a first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy includes: Based on the first identity type and the first behavior type, a first interactive voice with the aforementioned pronunciation timbre is generated for interaction.

10. The method according to claim 9, characterized in that, The step of generating a first interactive voice with the aforementioned timbre for interaction, based on the first identity type and the first behavior type, includes: Obtain the characteristic data of the pronunciation timbre; The feature data of the pronunciation timbre, the first identity type, and the first behavior type are input into the pre-trained sixth model to obtain the first interactive voice with the pronunciation timbre for interaction.

11. The method according to claim 9, characterized in that, The step of generating a first interactive voice with the aforementioned timbre for interaction, based on the first identity type and the first behavior type, includes: Obtain the seventh model corresponding to the pronunciation timbre; The first identity type and the first behavior type are input into the seventh model corresponding to the voice timbre to obtain a first interactive voice that is used for interaction and has the pronunciation timbre.

12. The method according to claim 1, characterized in that, Before determining the first voice interaction strategy that matches the first visitor, the method further includes: Determine the interviewee of the first visitor; The object information of the interviewees is determined, wherein the object information includes at least one of the following: the interviewee's vocal timbre, the interviewee's preference information, the interviewee's age, the interviewee's gender, and the number of interviewees; and The first voice interaction strategy for determining the match with the first visitor includes: Based on the object information, a first voice interaction strategy matching the first visiting object is determined.

13. The method according to claim 12, characterized in that, When the object information includes the preference information of the interviewee, determining a first voice interaction strategy matching the first interviewee based on the object information includes: Determine whether the preference information includes a voice interaction strategy triggered by the first visitor; If the preference information includes a voice interaction strategy triggered by the first visitor, the voice interaction strategy triggered by the first visitor is determined as a first voice interaction strategy that matches the first visitor. If the preference information does not include the voice interaction strategy triggered by the first visitor, a first voice interaction strategy matching the first visitor is generated based on the preference information.

14. The method according to claim 1, characterized in that, Before determining the first voice interaction strategy that matches the first visitor, the method further includes: Based on the first security video, determine the environmental information of the environment where the first visitor is located; and The first voice interaction strategy for determining the match with the first visitor includes: Based on at least one of the environmental information, the first identity type, and the first behavior type, a first voice interaction strategy matching the first visitor is determined.

15. The method according to claim 1, characterized in that, The step of determining the first identity type and first behavior type of the first visitor based on the first security video includes: The first security video is input into a multimodal big data model to determine the first identity type and first behavior type of the first visitor; and / or The first voice interaction strategy for determining the match with the first visitor includes: Using a multimodal large model, determine a first voice interaction strategy that matches the first visitor; and / or The step of generating a first interactive voice for interaction based on the first identity type and the first behavior type according to the first voice interaction strategy includes: The first voice interaction strategy, the first identity type, and the first behavior type are input into the multimodal large model to generate the first interactive voice for interaction.

16. The method according to claim 1, characterized in that, After the security agent plays the first interactive voice, the method further includes: Obtain the recorded video of the security intelligent agent playing the first interactive voice; Determine the feedback information from the preset object to the recorded video; Acquire a second security video containing images of a second visitor; Based on the second security video, determine the second identity type and second behavior type of the second visitor; Based on the feedback information, a second voice interaction strategy matching the second visitor is determined; According to the second voice interaction strategy, based on the second identity type and the second behavior type, a second interactive voice is generated for interacting with the second visitor. Control the security intelligent agent to play the second interactive voice.

17. The method according to any one of claims 1-16, characterized in that, The security smart device is a smart doorbell.

18. A voice interaction device for security scenarios, characterized in that, The device includes: The first acquisition unit is used to acquire a first security video containing an image of the first visitor. The first determining unit is used to determine the first identity type and the first behavior type of the first visitor based on the first security video. The second determining unit is used to determine a first voice interaction strategy that matches the first visitor. The first generation unit is configured to generate a first interactive voice for interaction based on the first identity type and the first behavior type, according to the first voice interaction strategy. The first control unit is used to control the security intelligent agent to play the first interactive voice.

19. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor is configured to execute a computer program stored in the memory, wherein when the computer program is executed, it implements the voice interaction method in the security scenario according to any one of claims 1-17.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the voice interaction method in the security scenario according to any one of claims 1-17.