Voice wake-up methods and apparatuses for camera unit, electronic device and storage medium

By receiving user voice in the camera unit and generating execution instructions using a large language model, the problems of low interaction efficiency and high false wake-up rate in existing technologies are solved, achieving a more efficient and accurate voice wake-up effect.

WO2026060899A1PCT designated stage Publication Date: 2026-03-26SHENZHEN DASHI SCI & TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

The existing technology for voice wake-up of camera units suffers from low interaction efficiency and high false wake-up rate, which affects user experience.

Method used

By receiving user voice sent from the interactive device, voice feature data is generated, and the data is processed using a large language model to generate execution instructions that match the user's intent. The camera unit is then directly invoked to capture images, avoiding the use of fixed voice wake-up words.

Benefits of technology

It improves the wake-up efficiency of the camera unit, reduces the false wake-up rate, and enhances the human-computer interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025079664_26032026_PF_FP_ABST
    Figure CN2025079664_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Voice wake-up methods and apparatuses for a camera unit, an electronic device and a storage medium. A method comprises: receiving a first user voice collected by an interaction device and, on the basis of the first user voice, obtaining corresponding voice feature data, the voice feature data representing voice content of the first user voice (S101); by means of a large language model, processing the voice feature data to generate a first execution instruction, instruction content of the first execution instruction matching a user intention that corresponds to the voice content (S102); and sending the first execution instruction to the interaction device to instruct a camera unit of the interaction device to shoot a target image (S103). The method does not need fixed voice wake-up words to trigger camera units to perform shooting, thereby reducing the false wake-up rate and improving human-computer interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Voice wake-up method and device of camera unit, electronic equipment and storage medium

[0001] The present application claims priority to the Chinese patent application No. 202411315114.5, filed on September 19, 2024, and entitled "Voice wake-up method and device of camera unit, electronic equipment and storage medium", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to intelligent terminal technology, in particular to a voice wake-up method and device of camera unit, electronic equipment and storage medium. BACKGROUND

[0003] With the development of intelligent terminal technology, various intelligent terminal devices can detect the external environment through the built-in camera unit, thereby realizing a more efficient interaction mode with the user.

[0004] In the prior art, due to the high running power consumption of the camera unit on the device, the camera unit is usually turned on on demand, for example, the user inputs a fixed voice wake-up word to trigger the camera unit to take pictures, thereby obtaining image data and performing subsequent processing to realize the target function.

[0005] However, the voice wake-up method of the camera unit in the prior art has the problems of low interaction efficiency and high false wake-up rate, which affects the user experience

[0006] SUMMARY

[0007] The embodiments of the present disclosure provide a voice wake-up method and device of camera unit, electronic equipment and storage medium to overcome the problems of low wake-up efficiency and high false wake-up rate.

[0008] In a first aspect, the embodiments of the present disclosure provide a voice wake-up method of camera unit, applied to a backend device, comprising:

[0009] receiving a first user voice sent by an interaction device, and obtaining corresponding voice feature data according to the first user voice, the voice feature data representing the voice content of the first user voice; processing the voice feature data through a large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches the user intent corresponding to the voice content; sending the first execution instruction to the interaction device to call the camera unit of the interaction device to take a target image.

[0010] In a second aspect, the embodiments of the present disclosure provide a voice wake-up method of camera unit, applied to an interaction device, comprising:

[0011] sending first user voice to a backend device to control the backend device to generate a first execution instruction, wherein the first execution instruction is generated based on a large language model, and instruction content of the first execution instruction matches a user intent corresponding to voice content of the first user voice;

[0012] receiving the first execution instruction sent by the backend device, and in response to the first execution instruction, calling a camera unit to capture a target image.

[0013] In a third aspect, the embodiments of the present disclosure provide a voice wake-up method of a camera unit, applied to an interactive device, comprising:

[0014] obtaining first user voice collected by the interactive device, and obtaining corresponding voice feature data according to the first user voice, wherein the voice feature data represents voice content of the first user voice;

[0015] generating a first execution instruction by processing the voice feature data through a large language model, wherein instruction content of the first execution instruction matches a user intent corresponding to the voice content;

[0016] calling a camera unit of the interactive device to capture a target image by executing the first execution instruction.

[0017] In a fourth aspect, the embodiments of the present disclosure provide a voice wake-up device of a camera unit, applied to a backend device, comprising:

[0018] a transceiving module, configured to receive first user voice sent by an interactive device;

[0019] a processing module, configured to obtain corresponding voice feature data according to the first user voice, wherein the voice feature data represents voice content of the first user voice;

[0020] a generating module, configured to generate a first execution instruction by processing the voice feature data through a large language model, wherein instruction content of the first execution instruction matches a user intent corresponding to the voice content;

[0021] The transceiving module is further configured to send the first execution instruction to the interactive device, so as to call a camera unit of the interactive device to capture a target image.

[0022] In a fifth aspect, the embodiments of the present disclosure provide a voice wake-up device of a camera unit, applied to an interactive device, comprising:

[0023] The sending module is configured to send the first user voice to a backend device to control the backend device to generate a first execution instruction, wherein the first execution instruction is generated based on a large language model, and an instruction content of the first execution instruction matches a user intent corresponding to a voice content of the first user voice.

[0024] The processing module is configured to receive the first execution instruction sent by the backend device, and in response to the first execution instruction, invoke the camera unit to capture a target image.

[0025] In a sixth aspect, the embodiments of the present disclosure provide a voice wake-up device of a camera unit, applied to an interactive device, and including:

[0026] The obtaining module is configured to obtain a first user voice collected by the interactive device, and obtain corresponding voice feature data according to the first user voice, wherein the voice feature data represents a voice content of the first user voice.

[0027] The generating module is configured to process the voice feature data by using a large language model to generate a first execution instruction, wherein an instruction content of the first execution instruction matches a user intent corresponding to the voice content.

[0028] The execution module is configured to invoke a camera unit of the interactive device to capture a target image by executing the first execution instruction.

[0029] In a seventh aspect, the embodiments of the present disclosure provide an electronic device, including a processor and a memory.

[0030] The memory stores computer execution instructions.

[0031] The processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the voice wake-up method of the camera unit as described in the first aspect and various possible designs of the first aspect, or executes the voice wake-up method of the camera unit as described in the second aspect and various possible designs of the second aspect, or executes the voice wake-up method of the camera unit as described in the third aspect and various possible designs of the third aspect.

[0032] In an eighth aspect, the embodiments of the present disclosure provide a computer readable storage medium, which stores computer execution instructions, and when a processor executes the computer execution instructions, the voice wake-up method of the camera unit as described in the first aspect and various possible designs of the first aspect is implemented, or the voice wake-up method of the camera unit as described in the second aspect and various possible designs of the second aspect is implemented, or the voice wake-up method of the camera unit as described in the third aspect and various possible designs of the third aspect is implemented.

[0033] In a ninth aspect, the embodiments of the present disclosure provide a computer program product, which comprises a computer program. When the computer program is executed by a processor, the voice wake-up method of the camera unit is implemented, as described in the first aspect and various possible designs of the first aspect. Alternatively, when the computer program is executed by the processor, the voice wake-up method of the camera unit is implemented, as described in the second aspect and various possible designs of the second aspect. Alternatively, when the computer program is executed by the processor, the voice wake-up method of the camera unit is implemented, as described in the third aspect and various possible designs of the third aspect.

[0034] The present application also provides a computer program. When the computer program is executed by a processor, the steps of the voice wake-up method of the camera unit are implemented.

[0035] The voice wake-up method of the camera unit, the device, the electronic equipment and the storage medium provided by the embodiments can receive a first user voice sent by an interactive device, and obtain corresponding voice feature data according to the first user voice. The voice feature data represents the voice content of the first user voice. The voice feature data is processed by a large language model to generate a first execution instruction. The instruction content of the first execution instruction matches a user intent corresponding to the voice content. The first execution instruction is sent to the interactive device to call a camera unit of the interactive device to shoot a target image. By converting the first user voice into voice feature data that can be processed by a large language model, and then processing the voice feature data based on a large voice model, a first execution instruction matching the user intent corresponding to the voice content is generated. The first execution instruction is sent to the interactive device to call the camera unit to shoot the target image. The fixed voice wake-up word is not used to trigger the camera unit to take a photo, the wake-up efficiency is improved, the false wake-up rate is reduced, and the human-computer interaction experience is improved. BRIEF DESCRIPTION OF DRAWINGS

[0036] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application. In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0037] FIG. 1 is a diagram of an application scenario of the voice wake-up method of the camera unit according to the embodiments of the present disclosure;

[0038] FIG. 2 is a flowchart of the voice wake-up method of the camera unit according to the embodiments of the present disclosure;

[0039] FIG. 3 is a flowchart of the specific implementation of step S101 in the embodiment shown in FIG. 2;

[0040] FIG. 4 is a flow chart of a specific implementation of step S103 in the embodiment shown in FIG. 2;

[0041] FIG. 5 is an interaction signaling diagram between a terminal device and an interactive device according to an embodiment of the present disclosure;

[0042] FIG. 6 is a flow chart of a voice wake-up method of a camera unit according to an embodiment of the present disclosure;

[0043] FIG. 7 is an interaction process diagram for generating a first execution instruction according to an embodiment of the present disclosure;

[0044] FIG. 8 is a flow chart of a voice wake-up method of a camera unit according to an embodiment of the present disclosure;

[0045] FIG. 9 is a flow chart of a voice wake-up method of a camera unit according to an embodiment of the present disclosure;

[0046] FIG. 10 is a signaling diagram of a voice wake-up method of a camera unit according to an embodiment of the present disclosure;

[0047] FIG. 11 is a flow chart of a voice wake-up method of a camera unit according to an embodiment of the present disclosure;

[0048] FIG. 12 is a structural block diagram of a voice wake-up device of a camera unit according to an embodiment of the present disclosure;

[0049] FIG. 13 is a structural block diagram of a voice wake-up device of a camera unit according to an embodiment of the present disclosure;

[0050] FIG. 14 is a structural block diagram of a voice wake-up device of a camera unit according to an embodiment of the present disclosure;

[0051] FIG. 15 is a structural diagram of an electronic device according to an embodiment of the present disclosure;

[0052] FIG. 16 is a hardware structural diagram of an electronic device according to an embodiment of the present disclosure;

[0053] The implementation, functional features and advantages of the present disclosure will be further described with reference to the embodiments and the accompanying drawings. The above-described drawings have shown specific embodiments of the present disclosure, and more detailed descriptions will be given hereinafter. These drawings and written descriptions are not intended to limit the scope of the present disclosure concept in any way, but to illustrate the present disclosure concept to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0054] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present disclosure.

[0055] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0056] The application scenarios of the embodiments of the present disclosure are explained as follows:

[0057] FIG. 1 is an application scenario diagram of the voice wake-up method of the camera unit provided by the embodiments of the present disclosure. The voice wake-up method of the camera unit provided by the embodiments of the present disclosure can be applied in an application scenario of multi-dimensional interaction among people, machines and environments, more specifically, can be applied in the process of triggering the interaction device to start the camera unit. The step of waking up the camera unit for photographing in the above application scenario can be independently executed by the interaction device, can be independently executed by the backend device, or can be jointly executed by the backend device and the interaction device. Among them, the backend device can be a terminal device, a server or an electronic device having similar functions, for example, a smart phone, a server corresponding to an application service end, a wearable device, etc. The interaction device can be a terminal device or an electronic device having similar functions, for example, a wearable device, more specifically, for example, a smart earphone, smart glasses, etc., and the interaction device is provided with a camera unit, for example, a camera, for collecting images. As shown in FIG. 1, the backend device is a smart phone, and the interaction device is a smart earphone. The user wears the smart earphone and establishes a connection with the smart earphone by operating an application program running on the smart phone. Then, the user opens the camera unit of the smart earphone and takes a target image by operating the keys of the smart earphone or inputting a voice instruction (the content of the voice instruction is, for example, “take a photo”) through the microphone unit of the smart earphone. After that, the target image is returned to the terminal device for processing, for example, saving the taken image or recognizing the content in the image and further returning the recognition result to the smart earphone in the form of voice for playing.

[0058] In some embodiments, the backend device can implement the voice wake-up method of the camera unit provided in the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be program-level commands, machine instructions, or software instructions. The computer program can be a native program in the operating system or a software module; it can be a local application program, i.e., a program that needs to be installed in the operating system to run, or it can be a small program embedded in any APP, i.e., a program running based on a browser environment. In summary, the above computer-executable instructions can be any form of instructions, and the above computer programs can be any form of application programs, modules, or plug-ins, and the specific implementation form can be configured as needed. Further, in some embodiments, the server can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud function, network service, middleware service, domain name service, security service, content delivery network (CDN), and basic cloud computing services such as big data and artificial intelligence platforms, etc., wherein the cloud service can be an interactive processing service for calling by a terminal device.

[0059] In the related art, a user triggers a camera unit to capture an image by speaking a fixed voice wake-up word. However, in the above application scenario, due to the triggering method by the voice wake-up word, the user needs to accurately speak the voice wake-up word alone to wake up the camera unit to take a photo; at the same time, in order to consider the wake-up efficiency, the wake-up word is usually simple, thereby causing the problem of mis-triggering the wake-up word when performing voice input in daily life, resulting in high false wake-up rate and the like.

[0060] The embodiments of the present disclosure provide a voice wake-up method of a camera unit to solve the above problems.

[0061] Referring to FIG. 2, FIG. 2 is a flowchart of a voice wake-up method of a camera unit provided by the embodiments of the present disclosure. The method of the present embodiment can be applied in a backend device, and the voice wake-up method of the camera unit comprises:

[0062] Step S101: receiving a first user voice collected by an interactive device, and obtaining corresponding voice feature data according to the first user voice, wherein the voice feature data represents the voice content of the first user voice.

[0063] With reference to the application scenario diagram shown in FIG. 1, the terminal device in the embodiment can be the backend device, and the interactive device can be the wearable device. More specifically, the terminal device is, for example, a smart phone, and the wearable device is, for example, a smart earphone. That is, in the embodiment, the terminal device is taken as the execution subject to introduce the voice wake-up method of the camera unit provided in the embodiment. Specifically, the terminal device and the interactive device can communicate through a local area wireless network, the Internet, or Bluetooth, which can be set as needed. After the terminal device and the interactive device create a communication connection, the terminal device can receive the data sent by the interactive device, specifically, the first user voice (corresponding voice data) collected by the interactive device. Further, on the side of the interactive device, the interactive device can collect the sound signal in the environment through the built-in microphone unit, and then generate the voice data corresponding to the first user voice, and send it to the terminal device after coding, compression, and other processing steps, so that the terminal device can receive the first user voice.

[0064] After the terminal device receives the first user voice sent by the interactive device through the wireless network, the terminal device processes the first user voice to generate corresponding voice feature data, which represents the voice content of the first user voice. Specifically, the first user voice is a kind of audio data, and the voice feature data describes the features of the voice content of the first user voice, which can be presented in the form of a feature matrix or a sequence. For example, the voice feature data can be data that can be processed by a subsequent large language model after one or more processing steps such as sampling and pooling (calculating the average of multiple data) of the first user voice.

[0065] In a possible implementation, as shown in FIG. 3, the specific implementation of step S101 includes:

[0066] Step S1011: After receiving the initial user voice sent by the interactive device, it is detected whether a pre-trigger instruction sent by the interactive device is received, wherein the pre-trigger instruction is generated when the initial user voice collected by the interactive device contains a preset target wake-up word.

[0067] Step S1012: If the pre-trigger instruction sent by the interactive device is received, the initial user voice is determined as the first user voice.

[0068] Step S1013: If the pre-trigger instruction is not received, the initial user voice is processed to generate the first user voice.

[0069] Step S1014: According to the first user voice, the corresponding voice feature data is obtained.

[0070] Exemplarily, in a possible implementation, after the terminal device establishes a communication connection, the interaction device maintains the communication channel and continuously sends the collected voice data, i.e., the initial user voice, to the terminal device, and meanwhile, the terminal device side synchronously receives the initial user voice. Therefore, in order to improve the execution efficiency of the terminal device side, after receiving the initial user voice, the terminal device first checks whether the pre-trigger instruction sent in advance by the interaction device is received, wherein the pre-trigger instruction is generated when the initial user voice collected by the interaction device contains a preset target wake-up word, i.e., after the interaction device collects the environmental sound signal to obtain the initial user voice, the initial user voice is subjected to an initial content detection to determine whether the target wake-up word is contained, and the target wake-up word is, for example, "take a photo". The specific implementation is that the audio data of the initial user voice is detected based on the audio features of the preset target wake-up word, so as to determine whether the target wake-up word is contained. Then, if the interaction device detects that the initial user voice contains the target wake-up word, a pre-trigger instruction is sent to the terminal device as prior knowledge for the terminal device to process the initial voice signal.

[0071] On the other hand, in one case, if the terminal device receives the pre-trigger instruction, it means that the initial user voice contains the target trigger word, i.e., there is a high probability of corresponding to the user's intention of waking up the camera unit for target image shooting; in this case, the initial user voice is determined as the first user voice for subsequent processing; and in another case, if the pre-trigger instruction is not received, i.e., there is a low probability of corresponding to the user's intention of waking up the camera unit for target image shooting, in this case, the terminal device further processes the initial user voice to generate the first user voice for subsequent processing. Specifically, for example, the initial user voice is subjected to target wake-up word detection, and after confirming that the target wake-up word is contained, it is taken as the first user voice, or the initial user voice is subjected to noise reduction processing, and then it is taken as the first user voice.

[0072] Finally, the first user voice obtained based on the above steps is converted into corresponding voice feature data.

[0073] In another possible implementation, the embodiment also includes:

[0074] After receiving the first user voice collected by the interaction device, it is detected whether the pre-trigger instruction sent by the interaction device is received, wherein the pre-trigger instruction is generated when the first user voice collected by the interaction device contains a preset target wake-up word.

[0075] Correspondingly, the specific implementation manner of obtaining the corresponding speech feature data according to the first user speech in step S101 includes: if the pre-trigger instruction sent by the interaction device is received, the corresponding speech feature data is obtained according to the first user speech.

[0076] Exemplarily, in the step of the embodiment, after receiving the first user speech collected by the interaction device, the back-end device first detects whether the pre-trigger instruction has been received. The pre-trigger instruction is generated on the side of the interaction device, and is equivalent to pre-processing of the first user speech when the interaction device detects that the target wake-up word is contained in the first user speech. Then, if the back-end device has received the pre-trigger instruction, the first execution instruction is directly generated. If the back-end device does not receive the pre-trigger instruction, the first user instruction is converted into speech feature data, and in the subsequent step, a large language model is called for processing to generate the first execution instruction.

[0077] As introduced in the above embodiment step, in the embodiment, the wake-up word detection can be performed on the collected user speech by the interaction device, or by the back-end device. The specific function or service used in the wake-up word detection can be deployed on the side of the interaction device, on the side of the external device, or on the side of other devices of a third party, such as a cloud service (server), without specific limitation.

[0078] In the step of the embodiment, by detecting whether the pre-trigger instruction sent by the interaction device is received, the step of obtaining the first user speech is further performed, which realizes the pre-processing of the initial user speech on the side of the interaction device, reduces the resource consumption of the terminal device, makes the generated speech feature data have better typicality, improves the accuracy of generating the first execution instruction by processing the speech feature data through the large language model, and makes the first execution instruction better match the user intent.

[0079] Step S102: processing the speech feature data through a large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches the user intent corresponding to the speech content.

[0080] Exemplarily, after generating the speech feature data, the terminal device further processes the speech feature data through the calling of the large language model to generate the first execution instruction. The instruction content of the first execution instruction matches the user intent corresponding to the speech content, and the specific content of the first execution instruction is determined based on the communication protocol of the terminal device and the interaction device. In a possible implementation manner, the specific implementation manner of step S102 includes:

[0081] Step S1021: performing intent recognition on the speech feature data through the large language model to obtain intent category information representing the user intent;

[0082] Step S1022: According to the intent category information, when the user intent is equivalent to or belongs to the target user intent, a first execution instruction is generated, wherein the target user intent includes a processing request for visual content.

[0083] Exemplarily, in a possible implementation, the large language model understands the user intent corresponding to the voice content by performing intent recognition on the input voice feature data, and obtains intent category information. The intent category information can represent different user intents in the form of feature sequences or category identifiers, such as “starting the dialing function of the terminal device”, “querying XX information stored in the terminal device”, “recognizing objects in the environment”, and the like. Then, the intent category information and the preset target user intent are classified and judged. The target user intent includes a processing request for visual content, that is, the target user intent needs to call the camera unit of the interactive device to perform image capturing. Therefore, when the intent category represented by the intent category information is equivalent to or belongs to the intent category represented by the target user intent, a first execution instruction is generated, that is, the user intent represented by the intent category information needs to call the camera unit of the interactive device to perform image capturing. For example, the user intent represented by the intent category information is “recognizing objects in the environment”, which belongs to the target user intent “processing request for visual content”, that is, a subset of the target user intent. The subset of the target user intent (user intent) can also include “navigation in the current environment”, “recognizing information on billboards”, and the like. The mapping relationship between the user intent represented by the intent category information and the target user intent can be determined by text similarity comparison, semantic distance comparison, and the like, which will not be described in detail. After determining that the user intent is equivalent to or belongs to the target user intent, an instruction for instructing the interactive device to start the camera unit to perform image capturing is generated, that is, the first execution instruction.

[0084] Further, wherein the large language model (LLM) refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. Illustratively, the input of the large language model is a feature matrix or a feature sequence, i.e. the data format possessed by the speech feature data, and the output is descriptive text containing the first execution instruction. The large language model can be deployed on a server or other electronic device outside the terminal device, and the terminal device remotely calls the large language model to process the speech feature data and generate the first execution instruction. In short, the above steps are to use the large language model to input the speech feature data to perform content recognition on the first user voice, wherein the large language model is a model obtained by pre-training using specific sample data, and its specific implementation principle, training method and use method (i.e. prompt word design) are not described here.

[0085] Step S103: Send the first execution instruction to the interactive device to call the camera unit of the interactive device to shoot the target image.

[0086] Illustratively, after generating the first execution instruction, the first execution instruction is sent to the interactive device. In one possible implementation, the first execution instruction contains a target field indicating whether to start the camera and / or whether to use the camera to take a photo. When the user's intention is determined to be equivalent to or attributed to the target user intention in the previous step, an instruction identifier representing the need to trigger the camera unit to shoot an image can be generated. Then, the instruction identifier is recorded in the target field, so that the first execution instruction can call the camera unit of the interactive device to shoot the target image based on the instruction identifier in the target field after being received and parsed by the interactive device according to the preset protocol.

[0087] Further, in one possible implementation, the process of sending the first execution instruction to the interactive device to call the camera unit of the interactive device to shoot the target image is divided into two stages, specifically, as shown in FIG. 4, the specific implementation of step S103 includes:

[0088] Step S1031: Send the first execution instruction to the interactive device to set the camera unit of the interactive device to an enabled state;

[0089] Step S1032: After receiving the second user voice sent by the interactive device, generate a third execution instruction according to the second user voice;

[0090] Step S1033: When detecting that the camera unit of the interactive device is in the enabled state, send the third execution instruction to the interactive device to call the camera unit of the interactive device to shoot the target image.

[0091] Exemplarily, in the embodiment step, after the terminal device sends the first execution instruction to the interactive device, the camera unit of the interactive device is first set to an enabled state, then the terminal device further interacts with the user, that is, receives a second user voice through the interactive device, generates a third execution instruction through the second user voice, and then combines the third execution instruction and the enabled state of the camera unit to call the camera unit of the interactive device to capture a target image. The second user voice can be a confirmation voice issued by the user, that is, the terminal device confirms the user's intention through the second user voice after recognizing the user's intention through the first user voice, and finally controls the camera unit of the interactive device to capture the target image through the third execution instruction, thereby further reducing the false triggering problem of the camera unit.

[0092] Fig. 5 is an interaction signaling diagram between a terminal device and an interactive device provided by an embodiment of the present disclosure. The above process will be introduced below in combination with Fig. 5. As shown in Fig. 5, exemplarily, first, the terminal device sends the first execution instruction to the interactive device after generating the first execution instruction. The interactive device (smart earphone) sets the working state of the camera unit to an enabled state after receiving the first execution instruction, and plays a first prompt voice to the user. The content of the first prompt voice is, for example, "Do you confirm to call the camera unit to take a photo?". Then, the user issues a second user voice according to the heard first prompt voice. The content of the second user voice is, for example, "Confirm to take a photo". The interactive device receives the second user voice and sends it to the terminal device for processing. The terminal device generates a third execution instruction for the user to control the interactive device to take a photo after receiving the second user voice, and detects the working state of the camera unit of the interactive device. If the camera unit of the interactive device is in the enabled state (Y path), the terminal device sends the third execution instruction to the interactive device to call the camera unit to capture a target image. On the other hand, if the camera unit of the interactive device is in a non-enabled state (N path), specifically, for example, in an occupied state, that is, the camera unit is occupied by other processes, the terminal device can wait and re-detect the working state of the camera unit after a preset time. Finally, when the interactive device captures a target image through the camera unit, the terminal device returns the target image to the terminal device, and the terminal device processes the target image based on subsequent functional steps.

[0093] It should be noted that in the embodiment, the backend device calls a large language model to generate the first execution instruction. The large language model can be deployed locally on the backend device or externally on the backend device, such as a cloud service container or an external server. The backend device accesses the external device to call the large language model to execute the above-mentioned related steps.

[0094] In this embodiment, the first user voice sent by the interaction device is received, and corresponding voice feature data is obtained according to the first user voice, the voice feature data representing the voice content of the first user voice; the voice feature data is processed by the large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches the user intent corresponding to the voice content; the first execution instruction is sent to the interaction device to call the camera unit of the interaction device to shoot the target image. By converting the first user voice into voice feature data that can be processed by the large language model, and then processing the voice feature data based on the large language model, a first execution instruction matching the user intent corresponding to the voice content is generated, and the first execution instruction is sent to the interaction device to call the camera unit to shoot the target image, without using a fixed voice wake-up word to trigger the camera unit to take a photo, improving the wake-up efficiency, reducing the false wake-up rate, and improving the human-computer interaction experience.

[0095] Referring to FIG. 6, FIG. 6 is a flowchart of a voice wake-up method of a camera unit according to an embodiment of the present disclosure. The embodiment further refines the wake-up process of the camera unit based on the embodiment shown in FIG. 2. The voice wake-up method of the camera unit includes the following steps:

[0096] Step S201: receiving a first user voice sent by an interaction device.

[0097] Step S202: detecting a target wake-up word in the first user voice.

[0098] Step S203: if the first user voice contains the target wake-up word, generating a first execution instruction corresponding to the target wake-up word.

[0099] Step S204: if the first user voice does not contain the target wake-up word, generating corresponding voice feature data according to the first user voice.

[0100] For example, in this embodiment, after the terminal device receives the first user voice sent by the interaction device, it first detects the target wake-up word in the first user voice, that is, it detects whether there is a specific wake-up word in the first user voice, thereby preliminarily judging the voice content of the first user voice. Compared with the step of calling the large language model for semantic recognition, this step has a smaller calculation amount, thereby reducing the resource consumption of the terminal device. Then, according to the detection result, different subsequent steps are executed.

[0101] Specifically, if the first user voice contains the target wake-up word, it is highly probable that it corresponds to the user's intention to wake up the camera unit to take a target image. In this case, the terminal device can directly generate the first execution instruction, and skip the step of calling the large model for semantic judgment, thereby greatly reducing the resource consumption of the terminal device, network resources and model resources, and improving the response speed. On the other hand, if the first user voice does not contain the target wake-up word, it needs to perform subsequent generation of speech feature data and call the large model for semantic judgment to understand the user's intention, and finally trigger the camera unit to take a picture. The specific implementation process of generating corresponding speech feature data from the first user voice has been introduced in the embodiment shown in FIG. 2, and will not be repeated here.

[0102] In the step of the present embodiment, two-stage detection is performed in combination with the wake-up word to realize a differentiated recognition scheme for the first user voice, thereby avoiding the problem of high resource consumption caused by processing through a large speech model, and improving the wake-up speed of the camera unit and the real-time performance of the image acquisition process.

[0103] Step S205: processing the speech feature data through the large language model to generate a first execution instruction and an image processing instruction.

[0104] Further, in the case where the first user voice does not contain the target wake-up word, the terminal device generates speech feature data, processes the speech feature data through the calling of the large language model, and generates a first execution instruction and an image processing instruction. The first execution instruction is an instruction for calling the camera unit of the interactive device to take a target image, and its specific generation process has been introduced in the previous embodiments and will not be repeated here. The image processing instruction is an instruction for processing the target image, which is generated or executed after the terminal device obtains the target image, thereby realizing the corresponding application function, such as recognizing the object in the target image, planning the navigation path based on the target image, etc. The specific implementation manner can be set as needed, which is not limited here. In the step of the present embodiment, the large language model processes the input data to output the first execution instruction and the image processing instruction. The first execution instruction and the image processing instruction can be recorded in the same text file or text segment output by the large language model, or presented in other text forms, which are not limited here.

[0105] Further, in one possible implementation manner, the present embodiment further includes:

[0106] Step S200: receiving image data taken by the interactive device, the image data being used to describe the shooting environment where the interactive device is located.

[0107] Correspondingly, the specific implementation manner of step S205 includes:

[0108] Step S205A: processing the voice feature data and the image data by the large language model to generate the first execution instruction and the image processing instruction.

[0109] Exemplarily, before obtaining the first user voice sent by the interaction device, the terminal device also receives image data sent by the interaction device, the image data being a picture taken by the interaction device to describe the shooting environment where the interaction device is located. In a possible implementation manner, the interaction device takes an image of the environment where the terminal device is located based on a fixed event interval or other trigger logic, and caches the image data (including one or more images taken by the interaction device) to send to the terminal device. The image data is used to describe the scene information of the current environment. After the terminal device processes the image data, the terminal device inputs the image data and the voice feature data corresponding to the received first user voice into the large language model. In this case, the large language model is a multi-modal model, which comprehensively identifies the user intent by combining the information in the image data and the voice feature data, and generates the first execution instruction matched with the user intent.

[0110] FIG. 7 is a schematic diagram of an interaction process for generating a first execution instruction according to an embodiment of the present disclosure. As shown in FIG. 7, first, the interaction device collects image data based on a preset time interval and sends the image data to the terminal device, for example, once every 1 minute. The terminal device caches the image data locally. Then, after receiving the first user voice sent by the interaction device, the terminal device generates corresponding voice feature data and inputs the voice feature data and the latest cached image data into the large language model deployed on an external server. After processing the voice feature data and the image data, the large language model generates a first execution instruction and returns the first execution instruction to the terminal device, thereby completing the generation process of the first execution instruction.

[0111] Optionally, before step S200, the embodiment further includes:

[0112] Step S200A: sending a fourth execution instruction to the interaction device, the fourth execution instruction being used to instruct the interaction device to take a reference image based on a target time interval.

[0113] Correspondingly, the specific implementation manner of step S200 includes: receiving a reference image sent by the interaction device based on a target time interval.

[0114] Exemplarily, the configuration of the parameters for the interactive device to collect the environmental image (reference image) can be achieved by an application program running on the terminal device. Specifically, the application program running in the terminal device sends a fourth execution instruction to the interactive device, and the fourth execution instruction configures the target time interval for the interactive device to collect the environmental image and the target resolution of the captured environmental image. Then, the interactive device generates the reference image based on the above configuration information and sends the reference image to the terminal device based on the target time interval, so as to achieve the sending purpose of the image data.

[0115] Further, in a possible implementation, the interactive device has at least two camera units located at different positions. For example, in the case of a smart earphone, one camera unit is arranged on each of the left earphone body and the right earphone body. The method further includes:

[0116] Step S200B: determining a target camera unit from the at least two camera units located at different positions according to the image data.

[0117] Correspondingly, the specific implementation of step S205A is that the large language model processes the voice feature data and the image data corresponding to the target camera unit to generate the first execution instruction matching the user intent of the voice content.

[0118] Exemplarily, the image data includes a pair of images captured by the camera units at different positions, for example, an image P1 captured by the camera unit on the left earphone body and an image P2 captured by the camera unit on the right earphone body. Therefore, the image data can be represented as

P1, P2

[0119] In a possible implementation, the camera units at different positions have different image shooting resolutions and / or shooting angles, that is, the resolutions and angles of the images shot by different camera units can be different, so as to realize different functions. In this case, in an implementation, a specific camera unit can be taken as a target camera unit according to parameters such as the resolution and shooting angle of image data, and then the image data shot by the target camera unit is taken as an input variable and the voice feature data are input into the large language model to perform intent recognition and generate a first execution instruction. Since the images shot by the camera units at different positions are screened, more target image data suitable for inputting into the large language model is obtained, for example, an image with lower resolution but larger shooting angle and wider field of view is taken as target image data to input into the large language model, so that the effective information amount is improved, and the first execution instruction generated by the large language model has higher accuracy.

[0120] Step S206: detecting a working state of the camera unit of the interactive device, wherein the working state at least includes a starting state or a closing state, and the starting state of the interactive device is triggered based on a second execution instruction, and the second execution instruction is generated before the first execution instruction.

[0121] Step S207: when the camera unit is in the starting state, the first execution instruction is sent to the interactive device to call the camera unit of the interactive device to shoot a target image.

[0122] Optionally, in this embodiment, after the terminal device generates the first execution instruction, it first detects the working state of the camera unit of the interactive device, wherein the working state at least includes a start state or a shutdown state. The start state can also be understood as the state of the camera unit after being started and initialized. When the camera unit is in the start state, it can directly capture images. When the camera unit is in the shutdown state, it needs to enter the start state (i.e., be initialized) first, and then it can capture images. On this basis, in this embodiment, the method can further include the step of sending a second execution instruction to the interactive device. For example, the second execution instruction can be similar to the first execution instruction. The second execution instruction is generated by the terminal device based on the historical voice sent by the interactive device after the terminal device performs semantic recognition or wake-up word recognition, and determines that the user has an intention to collect and process visual content. The generated instruction can also be an instruction generated by other triggering methods, such as an interactive operation on the terminal device side or a specific software control. The terminal device sends the second execution instruction to the interactive device to pre-trigger the camera unit of the interactive device, so that the camera unit is in the start state. Then, when the terminal device generates the first execution instruction and detects that the camera unit is in the start state, the terminal device sends the first execution instruction to the interactive device, so as to quickly call the camera unit and enable it to quickly capture the target image, thereby providing real-time performance of the captured target image. Further, the accuracy and real-time performance of subsequent content recognition of the target image are provided.

[0123] Step S208: performing content recognition on the target image in response to the image processing instruction to obtain a recognition result.

[0124] Step S209: generating an output voice based on the recognition result and sending the output voice to the interactive device to enable the interactive device to play the output voice or the backend device to play the output voice.

[0125] Further, after the terminal device obtains the target image, the terminal device performs further content recognition on the target image based on the image processing instruction obtained in the previous step to obtain a recognition result. Then, the terminal device generates an output voice based on the recognition result and sends the output voice to the interactive device to enable the interactive device to play the output voice. For example, the content of the first user voice is “XXX, what is this tree”, wherein “XXX” is a wake-up word. After the terminal device generates the target image corresponding to the first user voice and the image processing instruction based on the processing steps provided in the previous embodiment, the terminal device performs content recognition on the target image in response to the image processing instruction to generate a recognition result of the text “one sophora tree”. Then, the terminal device converts the text “one sophora tree” into an output voice and sends the output voice to the interactive device, such as a smart earphone, and plays it to the user. Thus, the recognition and interaction process between the user, the interactive device, and the environment is completed.

[0126] Or, in another possible implementation, the backend device directly plays the output voice described above.

[0127] In this embodiment, the implementation of step S201 is the same as that of step S101 in the embodiment shown in FIG. 2 of the present disclosure, and will not be repeated here.

[0128] Corresponding to the voice wake-up method for the camera unit of the backend device provided in the above embodiments, FIG. 8 is a third flowchart of the voice wake-up method for the camera unit provided in an embodiment of the present disclosure. The method of this embodiment can be applied in an interactive device. The voice wake-up method for the camera unit includes:

[0129] Step S301: sending a first user voice to a backend device to control the backend device to generate a first execution instruction, wherein the first execution instruction is generated based on a large language model, and the instruction content of the first execution instruction matches a user intent corresponding to the voice content of the first user voice;

[0130] Step S302: receiving the first execution instruction sent by the backend device, and in response to the first execution instruction, calling the camera unit to capture a target image.

[0131] Exemplarily, the backend device can be a terminal device or a server, such as a smart phone as introduced in the previous embodiments. The interactive device can be a wearable device, such as a smart earphone. In combination with the content of the embodiments corresponding to FIGS. 2-7, after the interactive device collects the first user voice, it sends the first user voice to the backend device. Based on the application protocol between the exchange device and the backend device, after the backend device receives the first user voice, it converts the first user voice into voice feature data, and then processes the voice feature data using a large language model to generate a first execution instruction matching the user intent. After that, the backend device sends the first execution instruction back to the interactive device. The interactive device calls the camera unit to capture a target image by responding to the first execution instruction.

[0132] The specific implementation of the above steps can refer to the introduction of the corresponding content in the embodiments corresponding to FIGS. 2-7, which will not be repeated here.

[0133] In this embodiment, the first user voice is sent to the backend device to control the backend device to generate a first execution instruction, wherein the first execution instruction is generated based on a large language model, and the instruction content of the first execution instruction matches a user intent corresponding to the voice content of the first user voice; the first execution instruction sent by the backend device is received, and the camera unit is called to capture a target image in response to the first execution instruction. By sending the first user voice to the backend device, the backend device converts the first user voice into a first execution instruction matching the user intent through the large language model, and then sends the first execution instruction to the interactive device. The interactive device responds to the first execution instruction to call the camera unit to capture the target image, without using a fixed voice wake-up word to trigger the camera unit to take a photo, thereby improving the wake-up efficiency, reducing the false wake-up rate, and improving the human-computer interaction experience.

[0134] FIG. 9 is a flowchart of a voice wake-up method of a camera unit according to an embodiment of the present disclosure. The voice wake-up method of the camera unit according to this embodiment is based on the embodiment shown in FIG. 8, and further refines the wake-up process of the camera unit. The voice wake-up method of the camera unit includes the following steps:

[0135] Step S300: Obtain image data, which is used to describe a shooting environment in which the interactive device is located.

[0136] Step S301: Obtain the first user voice by collecting an environmental sound signal.

[0137] Step S302: Detect a target wake-up word in the first user voice.

[0138] Step S303: If the target wake-up word is included in the first user voice, generate a pre-trigger instruction, and send the pre-trigger instruction and the first user voice to the backend device.

[0139] Step S304: If the target wake-up word is not included in the first user voice, send the first user voice and the image data to the backend device.

[0140] Optionally, the embodiment further includes the following steps:

[0141] Step S305: Receive a second execution instruction sent by the backend device, wherein the second execution instruction is generated by the large language model based on historical voice before the first user voice.

[0142] Step S306: In response to the second execution instruction, set the camera unit to an activated state.

[0143] Step S307: Receive a first execution instruction sent by the backend device.

[0144] Step S308: If the camera unit is in the activated state, call the camera unit to capture a target image in response to the first execution instruction.

[0145] Optionally, before step S300, further comprising

[0146] Step S300A: in response to the fourth execution instruction sent by the backend device, a reference image is captured based on a target time interval. Correspondingly, the specific implementation of step S300 includes: based on the reference image, image data is generated.

[0147] Exemplarily, the specific implementation of each step in this embodiment has been described in detail in the embodiments corresponding to FIGS. 2 to 7, which will not be repeated here.

[0148] FIG. 10 is a signaling diagram of a voice wake-up method of a camera unit provided by an embodiment of the present disclosure. The above embodiments will be described in more detail below in combination with FIG. 10. As shown in FIG. 10, exemplarily, first, the interactive device collects image data P1 and sends the image data P1 to the backend device, and the backend device caches the image data P1 locally. Then, the interactive device collects environment generated signals to generate a first user voice, and then detects the first user voice. If the first user voice contains a target wake-up word (Y path), a pre-trigger instruction and the first user voice are sent to the backend device. If the first user voice does not contain the target wake-up word (N path), only the pre-trigger instruction is sent to the backend device. The backend device receives the data sent by the interactive device, and if the data contains the pre-trigger instruction, a first execution instruction is directly generated. If the data does not contain the pre-trigger instruction, the first user voice is processed, for example, noise reduction, and then voice feature data is generated, and then a first execution instruction is generated based on the voice feature data and the image data P1, for example, a large language model. Then, the backend device sends the first execution instruction to the interactive device, and the interactive device receives the first execution instruction and sets the working state of the camera unit to the start state. Then, the interactive device collects and sends a second user voice to the backend device, and the terminal device receives the second user voice and generates a third execution instruction by using a large language model in combination with the image data P1, and sends the third execution instruction to the interactive device. The interactive device responds to the third execution instruction to capture a target image P2. Then, the interactive device sends the target image P2 to the backend device for processing, and the backend device processes the target image P2, for example, performs content recognition on the target image P2, and sends voice data corresponding to the recognition result to the interactive device for playing.

[0149] Further, in the above embodiments, the step of detecting the target wake-up word from the first user voice can be performed by the interactive device or the backend device. The specific function or service used for wake-up word detection can be deployed on the interactive device side, the external device side, or the third-party device side, such as a cloud service (server), without specific limitation.

[0150] FIG. 11 is a flowchart of a voice wake-up method of a camera unit according to an embodiment of the present disclosure. In one embodiment, the present disclosure further provides another voice wake-up method of a camera unit, which is applied to an interactive device. As shown in FIG. 11, the method includes the following steps.

[0151] In step S401, the first user voice collected by the interactive device is obtained, and corresponding voice feature data is obtained according to the first user voice, wherein the voice feature data represents the voice content of the first user voice.

[0152] In step S402, the voice feature data is processed by a large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches the user intent corresponding to the voice content.

[0153] In step S403, the camera unit of the interactive device is called to capture a target image by executing the first execution instruction.

[0154] In this embodiment, the first user voice collected by the interactive device is processed to generate voice feature data. Then, the first execution instruction is generated by calling the large language model to process the voice feature data, and then the camera unit is called to capture the target image by executing the first execution instruction. In the above process, the function or service used to process the first user voice collected by the interactive device to generate voice feature data can be deployed locally on the interactive device or externally, such as a cloud server. The interactive device can generate voice feature data by directly executing the corresponding local function, or by remotely calling external functions or services. On the other hand, the large language model can be deployed locally on the interactive device or externally, such as a cloud server, a container, or a backend device. When the large language model is deployed locally on the interactive device, the interactive device generates the corresponding first execution instruction by calling the local large language model. When the large language model is deployed externally on the interactive device, the interactive device sends the voice feature data to the external device to call the large language model to process the voice feature data. The above steps of generating voice feature data and generating the first execution instruction by processing the voice feature data through the large language model can be independent of each other, or can be implemented by the same local function or external function or external service, which is not limited here.

[0155] Further, in one possible implementation, after the interactive device obtains the first user voice, the present embodiment further includes the following steps:

[0156] In step S402, the voice feature data is processed by a large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches the user intent corresponding to the voice content.

[0157] According to the first user voice, corresponding voice feature data is obtained, including:

[0158] If the target wake-up word is not contained in the first user voice, the corresponding voice feature data is generated according to the first user voice.

[0159] In the steps of the embodiment, the interaction device performs wake-up word detection on the first user voice. If the target wake-up word is contained in the first user voice, a first execution instruction corresponding to the target wake-up word is directly generated to start the camera unit to capture the target image. If the target wake-up word is not contained in the first user voice, the subsequent steps of generating voice feature data and calling a large language model are continued. By detecting the target wake-up word, the preprocessing of the first user voice is realized, and according to the preprocessing result, the step of skipping calling the large model for semantic judgment is realized, thereby greatly reducing the resource consumption of the terminal device, network resources and model resources, and improving the response speed.

[0160] The specific function or service used by the interaction device when performing wake-up word detection on the first user voice can be deployed on the side of the interaction device, on the side of an external device, or on the side of other devices of a third party, such as a cloud service (server), without specific limitation.

[0161] Further, in a possible implementation manner, the embodiment further includes:

[0162] The working state of the camera unit of the interaction device is detected, wherein the working state at least includes a start state or a shutdown state, and the start state of the interaction device is triggered based on a second execution instruction generated before the first execution instruction; accordingly, the camera unit of the interaction device is called to capture the target image by executing the first execution instruction, including:

[0163] When the camera unit is in the start state, the first execution instruction is sent to the interaction device to call the camera unit of the interaction device to capture the target image.

[0164] In the steps of the embodiment, the interaction device detects the working state of the camera unit, and determines the subsequent execution steps according to the working state of the camera unit. Specifically, the interaction device is set to the start state after receiving the second execution instruction generated before the first execution instruction, and the first execution instruction is executed when the camera unit is in the start state. On the other hand, for example, when the camera unit is in the shutdown state or the occupied state, i.e., the camera unit is occupied by other processes, the working state of the camera unit can be detected again after a preset time. Finally, when the interaction device captures the target image through the camera unit, the target image is returned to the terminal device, and the terminal device processes the target image based on subsequent functional steps.

[0165] In another implementation manner, in the embodiment, when the photographing of the target image by the camera unit of the interactive device is invoked by executing the first execution instruction, the implementation manner comprises: setting the camera unit of the interactive device to an enabled state by executing the first execution instruction; collecting a second user voice, and generating a third execution instruction according to the second user voice; and executing the third execution instruction to invoke the camera unit of the interactive device to photograph the target image when it is detected that the camera unit of the interactive device is in the enabled state.

[0166] In the step of the embodiment, after the terminal device sends the first execution instruction to the interactive device, the camera unit of the interactive device is first set to an enabled state, and then the terminal device further interacts with the user, that is, the terminal device receives a second user voice through the interactive device, and then generates a third execution instruction through the second user voice. Then, the camera unit of the interactive device is invoked to photograph the target image in combination with the third execution instruction and the enabled state of the camera unit. The second user voice can be a confirmation voice issued by the user, that is, the terminal device confirms the user intention through the second user voice after recognizing the user intention through the first user voice. Finally, the camera unit of the interactive device is controlled to photograph the target image through the third execution instruction, so as to further reduce the false triggering problem of the camera unit.

[0167] Further, the embodiment further comprises: image data photographed, the image data being used to describe a photographing environment in which the interactive device is located; and generating the first execution instruction by processing the voice feature data through the large language model, comprising: generating the first execution instruction by processing the voice feature data and the image data through the large language model.

[0168] In the step of the embodiment, the interactive unit collects the first user information, and also periodically collects surrounding environment images, that is, image data, through the camera unit. Then, in the process of generating the first execution instruction, the voice feature data and the image data (corresponding image feature data) are processed through the large language model (multimodal model) to generate the first execution instruction. Since the information in the image data can be fully utilized, the accuracy of the first execution instruction can be further improved.

[0169] The specific implementation manners of the steps in the embodiment have been introduced in the previous embodiments, and can be referred to the introduction of the corresponding steps performed by the interactive device or the backend device in the previous embodiments, which will not be described herein again.

[0170] Corresponding to the voice wake-up method of the camera unit performed by the backend device in the above embodiments, FIG. 12 is a structural block diagram of a voice wake-up apparatus of a camera unit provided by an embodiment of the present disclosure. The method applied to the backend device introduced in the above embodiments can be performed by the voice wake-up apparatus of the camera unit. The apparatus can be implemented in a software and / or hardware manner, and the apparatus can be integrated in an electronic device with certain data processing functions. The electronic device can include but is not limited to a mobile terminal with large data processing capability, and a desktop computer, a supercomputer, and other fixed terminals with large data processing capability.

[0171] For ease of illustration, only parts related to the embodiments of the present disclosure are shown. Referring to FIG. 12, the voice wake-up apparatus 4 of the camera unit includes:

[0172] The transceiver module 41 is configured to receive the first user voice sent by the interactive device.

[0173] The processing module 42 is configured to obtain corresponding voice feature data according to the first user voice, the voice feature data representing the voice content of the first user voice.

[0174] The generation module 43 is configured to process the voice feature data through a large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches the user intent corresponding to the voice content.

[0175] The transceiver module 41 is further configured to send the first execution instruction to the interactive device to call the camera unit of the interactive device to shoot the target image.

[0176] According to one or more embodiments of the present disclosure, the generation module 43 is specifically configured to: perform intent recognition on the voice feature data through the large language model to obtain intent category information representing the user intent; and generate the first execution instruction according to the intent category information when the user intent is equivalent to or belongs to a target user intent, wherein the target user intent includes a processing request for visual content.

[0177] According to one or more embodiments of the present disclosure, before obtaining the corresponding voice feature data according to the first user voice, the generation module 43 is further configured to: detect a target wake-up word in the first user voice; if the first user voice contains the target wake-up word, generate a first execution instruction corresponding to the target wake-up word; and when obtaining the corresponding voice feature data according to the first user voice, the generation module 43 is specifically configured to: if the first user voice does not contain the target wake-up word, generate the corresponding voice feature data according to the first user voice.

[0178] According to one or more embodiments of the present disclosure, the transceiving module 41 is specifically configured to, when receiving the first user voice sent by the interactive device, detect whether a pre-trigger instruction sent by the interactive device is received after receiving the first user voice sent by the interactive device, wherein the pre-trigger instruction is generated when the interactive device collects the first user voice containing a preset target wake-up word; if the pre-trigger instruction sent by the interactive device is received, the first user voice is determined as the first user voice; if the pre-trigger instruction is not received, the first user voice is processed to generate the first user voice.

[0179] According to one or more embodiments of the present disclosure, the generating module 43 is further configured to detect a working state of a camera unit of the interactive device, wherein the working state at least includes a starting state or a closing state, and the starting state of the interactive device is triggered based on a second execution instruction generated by the large language model based on historical voice before the first user voice; and the transceiving module 41 is specifically configured to send the first execution instruction to the interactive device to call the camera unit of the interactive device to shoot the target image when the camera unit is in the starting state.

[0180] According to one or more embodiments of the present disclosure, the transceiving module 41 is specifically configured to send the first execution instruction to the interactive device to set the camera unit of the interactive device to the starting state; generate a third execution instruction based on the second user voice after receiving the second user voice sent by the interactive device; and send the third execution instruction to the interactive device to call the camera unit of the interactive device to shoot the target image when it is detected that the camera unit of the interactive device is in the starting state.

[0181] According to one or more embodiments of the present disclosure, the transceiving module 41 is further configured to receive image data shot by the interactive device, and the image data is used to describe a shooting environment where the interactive device is located; and the generating module 43 is specifically configured to process the voice feature data and the image data by the large language model to generate the first execution instruction.

[0182] According to one or more embodiments of the present disclosure, the transceiving module 41 is further configured to send a fourth execution instruction to the interactive device, and the fourth execution instruction is used to instruct the interactive device to shoot a reference image based on a target time interval; and the transceiving module 41 is specifically configured to receive the reference image sent by the interactive device based on the target time interval when receiving the image data shot by the interactive device.

[0183] According to one or more embodiments of the present disclosure, the interaction device has at least two camera units located at different positions, and the generation module 43 is further configured to determine, according to the image data, a target camera unit from the at least two camera units located at different positions; and when generating, by the large language model, the first execution instruction matching the user intent corresponding to the voice content, the generation module 43 is specifically configured to generate, by the large language model, the first execution instruction matching the user intent corresponding to the voice content by processing the voice feature data and the image data corresponding to the target camera unit.

[0184] According to one or more embodiments of the present disclosure, the generation module 43 is further configured to generate, by the large language model, an image processing instruction by processing the voice feature data; perform content recognition on the target image to obtain a recognition result in response to the image processing instruction; and generate an output voice based on the recognition result, and the transceiver module 41 is further configured to send the output voice to the interaction device to enable the interaction device to play the output voice.

[0185] The voice wake-up device 4 of the camera unit provided in this embodiment can execute the technical solutions of the method embodiments executed by the backend device as described above, and has similar implementation principles and technical effects, which will not be described here again in this embodiment.

[0186] Corresponding to the voice wake-up method of the camera unit executed by the interaction device in the above embodiments, FIG. 13 is a structural block diagram of a voice wake-up device of a camera unit according to an embodiment of the present disclosure. The method applied to the interaction device as described in the above embodiments can be executed by the voice wake-up device of the camera unit. The device can be implemented in the form of software and / or hardware, and the device can be integrated in an electronic device having a certain data processing function. The electronic device can include but is not limited to a mobile terminal with large data processing capability, and a desktop computer, a supercomputer and other fixed terminals with large data processing capability.

[0187] For ease of illustration, only parts related to the embodiments of the present disclosure are shown. Referring to FIG. 13, the voice wake-up device 5 of the camera unit includes:

[0188] The sending module 51 is configured to send a first user voice to a backend device to control the backend device to generate a first execution instruction, wherein the first execution instruction is generated based on a large language model, and the instruction content of the first execution instruction matches a user intent corresponding to a voice content of the first user voice;

[0189] The receiving module 52 is configured to receive the first execution instruction sent by the backend device;

[0190] The processing module 53 is configured to respond to the first execution instruction and call the camera unit to capture a target image.

[0191] According to one or more embodiments of the present disclosure, the processing module 53 is further configured to: obtain an initial user voice by collecting an environmental sound signal; detect a target wake-up word in the initial user voice; and the sending module 51 is specifically configured to: if the initial user voice contains the target wake-up word, generate a pre-trigger instruction and send the pre-trigger instruction and the initial user voice to the backend device; and if the initial user voice does not contain the target wake-up word, send the initial user voice to the backend device.

[0192] According to one or more embodiments of the present disclosure, before receiving the first execution instruction sent by the backend device, the receiving module 52 is further configured to: receive a second execution instruction sent by the backend device, the second execution instruction being generated by the large language model based on historical voice before the first user voice; and the processing module 53 is further configured to: in response to the second execution instruction, set the camera unit to an activated state; and when the processing module 53 invokes the camera unit to capture the target image in response to the first execution instruction, the processing module 53 is specifically configured to: if the camera unit is in the activated state, invoke the camera unit to capture the target image in response to the first execution instruction.

[0193] According to one or more embodiments of the present disclosure, the processing module 53 is further configured to: obtain image data, the image data being used to describe a shooting environment in which the interactive device is located; and when the sending module 51 sends the first user voice to the backend device to control the backend device to generate the first execution instruction, the sending module 51 is specifically configured to: send the first user voice and the image data to the backend device to control the backend device to generate the first execution instruction.

[0194] According to one or more embodiments of the present disclosure, the processing module 53 is further configured to: in response to a fourth execution instruction sent by the backend device, capture a reference image based on a target time interval; and when the processing module 53 obtains the image data, the processing module 53 is specifically configured to: generate the image data based on the reference image.

[0195] Corresponding to the voice wake-up method of the camera unit performed by the interactive device in the above embodiments, FIG. 14 is a structural block diagram of a voice wake-up device of a camera unit according to an embodiment of the present disclosure. The method applied to the interactive device introduced in the above embodiments can be performed by the voice wake-up device of the camera unit. The device can be implemented in a software and / or hardware manner, and the device can be integrated in an electronic device having a certain data processing function. The electronic device can include but is not limited to a mobile terminal with large data processing capability, and a desktop computer, a supercomputer, and other fixed terminals with large data processing capability.

[0196] For ease of illustration, only parts related to the embodiments of the present disclosure are shown. Referring to FIG. 14, the voice wake-up device 6 of the camera unit includes:

[0197] The acquisition module 61 is configured to acquire a first user voice collected by the interactive device, and obtain corresponding voice feature data according to the first user voice, wherein the voice feature data represents voice content of the first user voice.

[0198] The generation module 62 is configured to generate a first execution instruction by processing the voice feature data through a large language model, wherein the instruction content of the first execution instruction matches a user intent corresponding to the voice content.

[0199] The execution module 63 is configured to call a camera unit of the interactive device to shoot a target image by executing the first execution instruction.

[0200] According to one or more embodiments of the present disclosure, the generation module 62 is further configured to: detect a target wake-up word in the first user voice; and generate a first execution instruction corresponding to the target wake-up word if the target wake-up word is included in the first user voice.

[0201] The acquisition module 61 is specifically configured to: generate corresponding voice feature data according to the first user voice if the target wake-up word is not included in the first user voice.

[0202] According to one or more embodiments of the present disclosure, the generation module 62 is further configured to: detect a working state of the camera unit of the interactive device, wherein the working state at least includes a start state or a shutdown state, and the start state of the interactive device is triggered based on a second execution instruction, and the second execution instruction is generated before the first execution instruction; and the execution module 63 is specifically configured to: send the first execution instruction to the interactive device to call the camera unit of the interactive device to shoot the target image when the camera unit is in the start state.

[0203] According to one or more embodiments of the present disclosure, the execution module 63 is specifically configured to: set the camera unit of the interactive device to the start state by executing the first execution instruction; collect a second user voice, and generate a third execution instruction according to the second user voice; and execute the third execution instruction to call the camera unit of the interactive device to shoot the target image when it is detected that the camera unit of the interactive device is in the start state.

[0204] According to one or more embodiments of the present disclosure, the execution module 63 is further configured to: shoot image data, wherein the image data is used to describe a shooting environment where the interactive device is located; and the generation module 62 is specifically configured to: process the voice feature data and the image data through the large language model to generate the first execution instruction when generating the first execution instruction by processing the voice feature data through the large language model.

[0205] FIG. 15 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. As shown in FIG. 15, the electronic device 7 includes:

[0206] The processor 71 and the memory 72 are connected through the bus 73.

[0207] The memory 72 stores computer-executable instructions.

[0208] The processor 71 executes the computer-executable instructions stored in the memory 72 to implement the voice wake-up method of the camera unit in the embodiments shown in FIGS. 2-11.

[0209] Optionally, the processor 71 and the memory 72 are connected through the bus 73.

[0210] The related descriptions can be understood by referring to the related descriptions and effects of the steps in the embodiments corresponding to FIGS. 2-11, which will not be repeated here. The electronic device can be the backend device or the interactive device in the above embodiments.

[0211] The embodiments of the present disclosure provide a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the voice wake-up method of the camera unit in any of the embodiments corresponding to FIGS. 2-11 is implemented.

[0212] The embodiments of the present disclosure provide a computer program product, which includes a computer program. When the computer program is executed by a processor, the voice wake-up method of the camera unit in any of the embodiments corresponding to FIGS. 2-11 is implemented.

[0213] To implement the above embodiments, the embodiments of the present disclosure further provide an electronic device.

[0214] Referring to FIG. 16, a structural schematic diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure is shown. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a personal digital assistant (PDA), a tablet computer, a portable multimedia player (PMP), a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 16 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.

[0215] As shown in FIG. 16, the electronic device 900 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901 that can perform various appropriate actions and processes according to programs stored in a Read Only Memory (ROM) 902 or loaded into a Random Access Memory (RAM) 903 from a storage device 908. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An Input / Output (I / O) interface 905 is also connected to the bus 904.

[0216] Generally, the following devices can be connected to the I / O interface 905: input devices 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 907 including, for example, a Liquid Crystal Display (LCD), a speaker, a vibrator, etc.; storage devices 908 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 909. The communication devices 909 can allow the electronic device 900 to communicate wirelessly or wired with other devices to exchange data. Although FIG. 16 shows the electronic device 900 with various devices, it should be understood that all of the shown devices are not required to be implemented or possessed. More or less devices can be alternatively implemented or possessed.

[0217] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 909, or installed from the storage devices 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0218] It should be noted that the computer-readable medium in the above disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.

[0219] The computer-readable medium described above can be contained in the electronic device described above; or can exist separately and not be assembled into the electronic device.

[0220] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0221] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0222] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It is also noted that each block of the block diagrams and / or flow diagrams and combinations of blocks in the block diagrams and / or flow diagrams can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or by combinations of dedicated hardware and computer instructions.

[0223] The units or modules described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit or module does not constitute a limitation on the unit itself.

[0224] The functions described in the above description above can be performed by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0225] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0226] In a first aspect, according to one or more embodiments of the present disclosure, a voice wake-up method of a camera unit is provided, applied to a backend device, comprising:

[0227] receiving a first user voice sent by an interactive device, and obtaining corresponding voice feature data according to the first user voice, the voice feature data representing voice content of the first user voice; processing the voice feature data through a large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches a user intent corresponding to the voice content; and sending the first execution instruction to the interactive device to call a camera unit of the interactive device to shoot a target image.

[0228] According to one or more embodiments of the present disclosure, the processing of the voice feature data through the large language model to generate the first execution instruction comprises: performing intent recognition on the voice feature data through the large language model to obtain intent category information representing the user intent; and according to the intent category information, when the user intent is equivalent to or belongs to a target user intent, generating the first execution instruction, wherein the target user intent includes a processing request for visual content.

[0229] According to one or more embodiments of the present disclosure, before the obtaining of the corresponding voice feature data according to the first user voice, further comprising: detecting a target wake-up word in the first user voice; if the first user voice contains the target wake-up word, generating a first execution instruction corresponding to the target wake-up word; and the obtaining of the corresponding voice feature data according to the first user voice comprises: if the first user voice does not contain the target wake-up word, generating the corresponding voice feature data according to the first user voice.

[0230] According to one or more embodiments of the present disclosure, the receiving the first user voice sent by the interaction device comprises: after receiving the initial user voice sent by the interaction device, detecting whether a pre-trigger instruction sent by the interaction device is received, wherein the pre-trigger instruction is generated when the initial user voice collected by the interaction device contains a preset target wake-up word; if the pre-trigger instruction sent by the interaction device is received, the initial user voice is determined as the first user voice; if the pre-trigger instruction is not received, the initial user voice is processed to generate the first user voice.

[0231] According to one or more embodiments of the present disclosure, the method further comprises: detecting a working state of a camera unit of the interaction device, wherein the working state at least comprises a starting state or a closing state, and the starting state of the interaction device is triggered based on a second execution instruction generated by the large language model based on historical voice before the first user voice; and the sending the first execution instruction to the interaction device to call the camera unit of the interaction device to shoot a target image comprises: when the camera unit is in the starting state, the first execution instruction is sent to the interaction device to call the camera unit of the interaction device to shoot a target image.

[0232] According to one or more embodiments of the present disclosure, the sending the first execution instruction to the interaction device to call the camera unit of the interaction device to shoot a target image comprises: sending the first execution instruction to the interaction device to set the camera unit of the interaction device to the starting state; after receiving a second user voice sent by the interaction device, generating a third execution instruction according to the second user voice; and when it is detected that the camera unit of the interaction device is in the starting state, sending the third execution instruction to the interaction device to call the camera unit of the interaction device to shoot a target image.

[0233] According to one or more embodiments of the present disclosure, the method further comprises: receiving image data shot by the interaction device, the image data being used to describe a shooting environment where the interaction device is located; and the processing the voice feature data by the large language model to generate the first execution instruction comprises: processing the voice feature data and the image data by the large language model to generate the first execution instruction.

[0234] According to one or more embodiments of the present disclosure, the method further comprises: sending a fourth execution instruction to the interaction device, the fourth execution instruction being used to instruct the interaction device to shoot a reference image based on a target time interval; and the receiving the image data shot by the interaction device comprises: receiving the reference image sent by the interaction device based on the target time interval.

[0235] According to one or more embodiments of the present disclosure, the interaction device has at least two camera units located at different positions, and the method further comprises: determining a target camera unit from the at least two camera units located at different positions according to the image data; and generating, by the large language model, the first execution instruction matching the user intent corresponding to the voice content by processing the voice feature data and the image data corresponding to the target camera unit.

[0236] According to one or more embodiments of the present disclosure, the method further comprises: generating, by the large language model, an image processing instruction by processing the voice feature data; performing content recognition on the target image to obtain a recognition result in response to the image processing instruction; generating an output voice based on the recognition result and sending the output voice to the interaction device to make the interaction device play the output voice.

[0237] In a second aspect, according to one or more embodiments of the present disclosure, a voice wake-up method of a camera unit is provided, comprising:

[0238] sending the first user voice to a backend device to control the backend device to generate a first execution instruction, wherein the first execution instruction is generated based on a large language model, and the instruction content of the first execution instruction matches a user intent corresponding to voice content of the first user voice; receiving the first execution instruction sent by the backend device, and calling a camera unit to capture a target image in response to the first execution instruction.

[0239] According to one or more embodiments of the present disclosure, the method further comprises: obtaining an initial user voice by collecting an environmental sound signal; detecting a target wake-up word in the initial user voice; and sending the first user voice to the backend device, comprising: if the initial user voice contains the target wake-up word, generating a pre-trigger instruction and sending the pre-trigger instruction and the initial user voice to the backend device; and if the initial user voice does not contain the target wake-up word, sending the initial user voice to the backend device.

[0240] According to one or more embodiments of the present disclosure, before the receiving the first execution instruction sent by the backend device, the method further comprises: receiving a second execution instruction sent by the backend device, the second execution instruction being generated by the large language model based on historical speech before the first user speech; in response to the second execution instruction, setting the camera unit to an activated state; and the calling the camera unit to capture a target image in response to the first execution instruction comprises: if the camera unit is in the activated state, calling the camera unit to capture the target image in response to the first execution instruction.

[0241] According to one or more embodiments of the present disclosure, the method further comprises: obtaining image data, the image data being used to describe a shooting environment in which the interactive device is located; and the sending the first user speech to the backend device to control the backend device to generate a first execution instruction comprises: sending the first user speech and the image data to the backend device to control the backend device to generate the first execution instruction.

[0242] According to one or more embodiments of the present disclosure, the method further comprises: in response to a fourth execution instruction sent by the backend device, capturing a reference image based on a target time interval; and the obtaining the image data comprises: generating the image data based on the reference image.

[0243] In a third aspect, the present disclosure provides a voice wake-up method of a camera unit, applied to an interactive device, comprising:

[0244] obtaining first user speech collected by the interactive device, and obtaining corresponding speech feature data according to the first user speech, the speech feature data representing speech content of the first user speech;

[0245] processing the speech feature data by a large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches a user intent corresponding to the speech content;

[0246] calling a camera unit of the interactive device to capture a target image by executing the first execution instruction.

[0247] In a fourth aspect, according to one or more embodiments of the present disclosure, a voice wake-up device of a camera unit is provided, applied to a backend device, comprising:

[0248] a transceiving module, configured to receive first user speech sent by an interactive device;

[0249] a processing module, configured to obtain corresponding speech feature data according to the first user speech, the speech feature data representing speech content of the first user speech;

[0250] The generating module is configured to generate a first execution instruction by processing the voice feature data through the large language model, wherein the instruction content of the first execution instruction matches the user intent corresponding to the voice content.

[0251] The transceiving module is further configured to send the first execution instruction to the interactive device to call the camera unit of the interactive device to capture the target image.

[0252] According to one or more embodiments of the present disclosure, the generating module is specifically configured to: perform intent recognition on the voice feature data through the large language model to obtain intent category information representing the user intent; and generate the first execution instruction when the user intent is equivalent to or belongs to a target user intent according to the intent category information, wherein the target user intent includes a processing request for visual content.

[0253] According to one or more embodiments of the present disclosure, before obtaining the corresponding voice feature data according to the first user voice, the generating module is further configured to: detect a target wake-up word in the first user voice; and generate the first execution instruction corresponding to the target wake-up word if the target wake-up word is included in the first user voice; and when obtaining the corresponding voice feature data according to the first user voice, the generating module is specifically configured to: generate the corresponding voice feature data according to the first user voice if the target wake-up word is not included in the first user voice.

[0254] According to one or more embodiments of the present disclosure, when receiving the first user voice sent by the interactive device, the transceiving module is specifically configured to: after receiving the initial user voice sent by the interactive device, detect whether a pre-trigger instruction sent by the interactive device is received, wherein the pre-trigger instruction is generated when the initial user voice collected by the interactive device includes a preset target wake-up word; and if the pre-trigger instruction sent by the interactive device is received, the initial user voice is determined as the first user voice; and if the pre-trigger instruction is not received, the initial user voice is processed to generate the first user voice.

[0255] According to one or more embodiments of the present disclosure, the generating module is further configured to: detect a working state of the camera unit of the interactive device, wherein the working state at least includes a start state or a shutdown state, and the start state of the interactive device is triggered based on a second execution instruction, and the second execution instruction is generated by the large language model based on historical voice before the first user voice; and the transceiving module is specifically configured to: when the camera unit is in the start state, send the first execution instruction to the interactive device to call the camera unit of the interactive device to capture the target image.

[0256] According to one or more embodiments of the present disclosure, the transceiving module is specifically configured to: send a first execution instruction to the interactive device, so as to set a camera unit of the interactive device to an activated state; after receiving a second user voice sent by the interactive device, generate a third execution instruction according to the second user voice; and when detecting that the camera unit of the interactive device is in the activated state, send the third execution instruction to the interactive device, so as to call the camera unit of the interactive device to capture a target image.

[0257] According to one or more embodiments of the present disclosure, the transceiving module is further configured to: receive image data captured by the interactive device, the image data being used to describe a shooting environment in which the interactive device is located; and the generating module is specifically configured to: process the voice feature data and the image data by using the large language model, and generate the first execution instruction.

[0258] According to one or more embodiments of the present disclosure, the transceiving module is further configured to: send a fourth execution instruction to the interactive device, the fourth execution instruction being used to instruct the interactive device to capture a reference image based on a target time interval; and when receiving the image data captured by the interactive device, the transceiving module is specifically configured to: receive the reference image sent by the interactive device based on the target time interval.

[0259] According to one or more embodiments of the present disclosure, the interactive device has at least two camera units located at different positions, and the generating module is further configured to: determine a target camera unit from the at least two camera units located at different positions according to the image data; and when generating the first execution instruction matching the user intent corresponding to the voice content by processing the voice feature data and the image data by using the large language model, the generating module is specifically configured to: generate the first execution instruction matching the user intent corresponding to the voice content by processing the voice feature data and image data corresponding to the target camera unit by using the large language model.

[0260] According to one or more embodiments of the present disclosure, the generating module is further configured to: process the voice feature data by using the large language model, and generate an image processing instruction; perform content recognition on the target image in response to the image processing instruction, and obtain a recognition result; and generate an output voice based on the recognition result, and the transceiving module is further configured to: send the output voice to the interactive device, so that the interactive device plays the output voice.

[0261] In a fifth aspect, according to one or more embodiments of the present disclosure, a voice wake-up device of a camera unit is provided, which comprises:

[0262] The sending module is configured to send a first user voice to a backend device, so as to control the backend device to generate a first execution instruction, wherein the first execution instruction is generated based on a large language model, and an instruction content of the first execution instruction matches a user intent corresponding to a voice content of the first user voice;

[0263] The receiving module is configured to receive a first execution instruction sent by the backend device.

[0264] The processing module is configured to, in response to the first execution instruction, invoke the camera unit to capture a target image.

[0265] According to one or more embodiments of the present disclosure, the processing module is further configured to: obtain an initial user voice by collecting an environmental sound signal; detect a target wake-up word in the initial user voice; and the sending module is specifically configured to: if the initial user voice contains the target wake-up word, generate a pre-trigger instruction and send the pre-trigger instruction and the initial user voice to the backend device; and if the initial user voice does not contain the target wake-up word, send the initial user voice to the backend device.

[0266] According to one or more embodiments of the present disclosure, before receiving the first execution instruction sent by the backend device, the receiving module is further configured to receive a second execution instruction sent by the backend device, the second execution instruction being generated by a large language model based on historical voice before the first user voice; the processing module is further configured to: in response to the second execution instruction, set the camera unit to an activated state; and when the processing module invokes the camera unit to capture the target image in response to the first execution instruction, the processing module is specifically configured to: if the camera unit is in the activated state, invoke the camera unit to capture the target image in response to the first execution instruction.

[0267] According to one or more embodiments of the present disclosure, the processing module is further configured to: obtain image data, the image data being used to describe a shooting environment in which the interactive device is located; and when the sending module sends the first user voice to the backend device to control the backend device to generate the first execution instruction, the sending module is specifically configured to send the first user voice and the image data to the backend device to control the backend device to generate the first execution instruction.

[0268] According to one or more embodiments of the present disclosure, the processing module is further configured to: in response to a fourth execution instruction sent by the backend device, capture a reference image based on a target time interval; and when the processing module obtains the image data, the processing module is specifically configured to: generate the image data based on the reference image.

[0269] In a fifth aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one processor and a memory;

[0270] The memory stores computer execution instructions;

[0271] The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the voice wake-up method of the camera unit as described in the first aspect and various possible designs of the first aspect, or executes the voice wake-up method of the camera unit as described in the second aspect and various possible designs of the second aspect.

[0272] In a sixth aspect, the embodiments of the present disclosure provide a voice wake-up device of a camera unit, applied to an interactive device, comprising:

[0273] an acquisition module, configured to acquire a first user voice collected by the interactive device, and obtain corresponding voice feature data according to the first user voice, the voice feature data representing voice content of the first user voice;

[0274] a generation module, configured to process the voice feature data by a large language model to generate a first execution instruction, wherein an instruction content of the first execution instruction matches a user intent corresponding to the voice content;

[0275] an execution module, configured to call a camera unit of the interactive device to shoot a target image by executing the first execution instruction.

[0276] In a seventh aspect, the embodiments of the present disclosure provide an electronic device, comprising a processor and a memory;

[0277] the memory stores computer execution instructions;

[0278] the processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the voice wake-up method of the camera unit as described in the first aspect and various possible designs of the first aspect; or executes the voice wake-up method of the camera unit as described in the second aspect and various possible designs of the second aspect; or executes the voice wake-up method of the camera unit as described in the third aspect and various possible designs of the third aspect.

[0279] In an eighth aspect, the embodiments of the present disclosure provide a computer readable storage medium, wherein the computer readable storage medium stores computer execution instructions, when the processor executes the computer execution instructions, the voice wake-up method of the camera unit as described in the first aspect and various possible designs of the first aspect is realized; or the voice wake-up method of the camera unit as described in the second aspect and various possible designs of the second aspect is realized; or the voice wake-up method of the camera unit as described in the third aspect and various possible designs of the third aspect is realized.

[0280] In a ninth aspect, the embodiments of the present disclosure provide a computer program product, comprising a computer program, when the computer program is executed by the processor, the voice wake-up method of the camera unit as described in the first aspect and various possible designs of the first aspect is realized; or the voice wake-up method of the camera unit as described in the second aspect and various possible designs of the second aspect is realized; or the voice wake-up method of the camera unit as described in the third aspect and various possible designs of the third aspect is realized.

[0281] The above description merely illustrates the preferred embodiments of the disclosure and a principle for applying the technologies. It is understood by those skilled in the art that the disclosed scope of the disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical solutions formed by the mutual replacement of the above-described features and the technical features with similar functions disclosed in the disclosure (but not limited to) can be formed.

[0282] Further, although operations are depicted in a particular, sequential order, this should not be understood as requiring or implying that the operations are performed in the order illustrated or sequentially. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are included for the purpose of providing a thorough disclosure, these should not be construed as limitations on the scope of the disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0283] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A voice wake-up method for a camera unit, characterized in that, Applied to a backend device, comprising: Receiving a first user voice collected by an interactive device, and obtaining corresponding voice feature data according to the first user voice, the voice feature data representing the voice content of the first user voice; Processing the voice feature data through a large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches the user intent corresponding to the voice content; Sending the first execution instruction to the interactive device to call the camera unit of the interactive device to shoot a target image.

2. The method of claim 1, wherein, The processing of the voice feature data through a large language model to generate a first execution instruction comprises: Performing intent recognition on the voice feature data through a large language model to obtain intent category information representing the user intent; According to the intent category information, when the user intent is equivalent to or belongs to a target user intent, the first execution instruction is generated, wherein the target user intent includes a processing request for visual content.

3. The method according to any of claims 1-2, characterized in that, Before the corresponding voice feature data is obtained according to the first user voice, it further comprises: Detecting a target wake-up word in the first user voice; If the first user voice contains the target wake-up word, a first execution instruction corresponding to the target wake-up word is generated; If the first user voice does not contain a target wake-up word, corresponding voice feature data is generated according to the first user voice. The method further comprises:

4. The method according to any one of claims 1 to 3, characterized in that, After receiving the first user voice collected by the interactive device, it is detected whether a pre-trigger instruction sent by the interactive device is received, wherein the pre-trigger instruction is generated when the first user voice collected by the interactive device contains a preset target wake-up word; If the first user voice contains the target wake-up word, a first execution instruction corresponding to the target wake-up word is generated; The method further comprises: Detecting the working state of the camera unit of the interactive device, wherein the working state at least includes a start state or a shutdown state, and the start state of the interactive device is triggered based on a second execution instruction, and the second execution instruction is generated before the first execution instruction; 5. The method according to any one of claims 1 to 4, characterized in that, When the camera unit is in the start state, the first execution instruction is sent to the interactive device to call the camera unit of the interactive device to shoot a target image. The first execution instruction is sent to the interactive device to set the camera unit of the interactive device to a start state; After receiving the second user voice sent by the interactive device, a third execution instruction is generated according to the second user voice; ​ 6. The method according to any one of claims 1 to 5, characterized in that, ​ ​ ​ In a case where it is detected that the camera unit of the interaction device is in the starting state, the third execution instruction is sent to the interaction device to call the camera unit of the interaction device to capture a target image.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: receiving image data captured by the interaction device, the image data being used to describe a capturing environment in which the interaction device is located; The processing of the speech feature data by the large language model includes: processing the speech feature data and the image data by the large language model to generate the first execution instruction.

8. The method of claim 7, wherein, Further includes: sending a fourth execution instruction to the interaction device, the fourth execution instruction being used to instruct the interaction device to capture a reference image based on a target time interval; The receiving of the image data captured by the interaction device includes: receiving the reference image sent by the interaction device based on the target time interval.

9. The method of claim 8, wherein, The interaction device has at least two camera units located at different positions, and the method further includes: determining a target camera unit from the at least two camera units located at different positions according to the image data; The processing of the speech feature data and the image data by the large language model to generate the first execution instruction matching the user intent corresponding to the speech content includes: processing the speech feature data and the image data corresponding to the target camera unit by the large language model to generate the first execution instruction matching the user intent corresponding to the speech content.

10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: processing the speech feature data by the large language model to generate an image processing instruction; performing content recognition on the target image in response to the image processing instruction to obtain a recognition result; generating an output speech based on the recognition result and sending the output speech to the interaction device to make the interaction device play the output speech, or making the backend device play the output speech.

11. The method according to any one of claims 1 to 10, characterized in that, The interaction device is configured to process the speech feature data by the large language model; or The interaction device is in wireless communication with a backend device, and the backend device is configured to process the speech feature data by the large language model.

12. The method of claims 1-11, wherein, The interaction device is configured to detect a target wake-up word in the first user speech; or The interaction device is in wireless communication with a backend device, and the backend device is configured to detect a target wake-up word in the first user speech.

13. A voice wake-up method of a camera unit, characterized by, Applied to an interaction device, including: sending a first user speech to a backend device to make the backend device generate a first execution instruction based on the first user speech, wherein the first execution instruction is generated based on a large language model, and the instruction content of the first execution instruction matches a user intent corresponding to the speech content of the first user speech; receiving the first execution instruction sent by the backend device and calling a camera unit of the interaction device to capture a target image in response to the first execution instruction.

14. The method of claim 13, wherein, The method further includes: obtaining a first user speech by collecting an environmental sound signal; detecting a target wake-up word in the first user speech; if the target wake-up word is included in the first user speech, generating a first execution instruction corresponding to the target wake-up word; If the target wake-up word is not included in the first user voice, the first user voice is sent to the backend device.

15. The method according to any of claims 13-14, characterized by, Before receiving the first execution instruction sent by the backend device, the method further includes: In response to the second execution instruction, the camera unit is set to an activated state. The response to the first execution instruction includes: If the camera unit is in the activated state, the camera unit is called to capture the target image in response to the first execution instruction.

16. The method according to any one of claims 13-15, characterized in that, The method further includes: Obtaining image data, the image data being used to describe a shooting environment in which the interactive device is located; The first user voice is sent to the backend device to control the backend device to generate a first execution instruction, including: The first user voice and the image data are sent to the backend device to control the backend device to generate a first execution instruction.

17. The method of claim 16, wherein, Further including: Capturing a reference image based on a target time interval; The image data is obtained, including: The image data is generated based on the reference image.

18. A voice wake-up method of a camera unit, characterized by, Applied to an interactive device, including: Obtaining a first user voice collected by the interactive device, and obtaining corresponding voice feature data according to the first user voice, the voice feature data representing the voice content of the first user voice; A first execution instruction is generated by processing the voice feature data through a large language model, wherein the instruction content of the first execution instruction matches the user intent corresponding to the voice content; A target image is captured by calling the camera unit of the interactive device through the execution of the first execution instruction.

19. The method of claim 18, wherein, The method further includes: Detecting a target wake-up word in the first user voice; If the target wake-up word is included in the first user voice, a first execution instruction corresponding to the target wake-up word is generated; The corresponding voice feature data is obtained according to the first user voice, including: If the target wake-up word is not included in the first user voice, the corresponding voice feature data is generated according to the first user voice.

20. The method of any one of claims 18-19, wherein, The method further includes: Detecting the working state of the camera unit of the interactive device, wherein the working state includes at least an activated state or a closed state, and the activated state of the interactive device is triggered based on a second execution instruction, and the second execution instruction is generated before the first execution instruction; The camera unit of the interactive device is called to capture a target image by executing the first execution instruction, including: When the camera unit is in the activated state, the first execution instruction is sent to the interactive device to call the camera unit of the interactive device to capture a target image.

21. The method according to any one of claims 18-20, characterized by, The camera unit of the interactive device is called to capture a target image by executing the first execution instruction, including: The camera unit of the interactive device is set to an activated state by executing the first execution instruction; A second user voice is collected, and a third execution instruction is generated according to the second user voice; When it is detected that the camera unit of the interactive device is in the activated state, the third execution instruction is executed to call the camera unit of the interactive device to capture a target image.

22. The method according to any one of claims 18-21, characterized by, The method further includes: Image data captured by the camera unit, the image data being used to describe a shooting environment in which the interactive device is located; The processing of the voice feature data by the large language model includes: The processing of the voice feature data and the image data by the large language model generates the first execution instruction.

23. A voice wake-up apparatus for a camera unit, characterized by Applied to a backend device, comprising: A transceiver module configured to receive a first user voice sent by an interactive device; A processing module configured to obtain corresponding voice feature data from the first user voice, the voice feature data representing a voice content of the first user voice; A generation module configured to process the voice feature data by a large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches a user intent corresponding to the voice content; The transceiver module is further configured to send the first execution instruction to the interactive device to call a camera unit of the interactive device to capture a target image.

24. A voice wake-up apparatus for a camera unit, characterized by Applied to an interactive device, comprising: A sending module configured to send a first user voice to a backend device to control the backend device to generate a first execution instruction, wherein the first execution instruction is generated based on a large language model, and the instruction content of the first execution instruction matches a user intent corresponding to a voice content of the first user voice; A receiving module configured to receive the first execution instruction sent by the backend device; A processing module configured to call a camera unit to capture a target image in response to the first execution instruction.

25. A voice wake-up device for a camera unit, characterized in that, Applied to an interactive device, comprising: An acquisition module configured to acquire a first user voice collected by an interactive device and obtain corresponding voice feature data from the first user voice, the voice feature data representing a voice content of the first user voice; A generation module configured to process the voice feature data by a large language model to generate a first execution instruction, wherein the instruction content of the first execution instruction matches a user intent corresponding to the voice content; An execution module configured to call a camera unit of the interactive device to capture a target image by executing the first execution instruction.

26. An electronic device, comprising: Comprising: A processor and a memory; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the voice wake-up method of the camera unit according to any one of claims 1 to 22.

27. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and when the processor executes the computer execution instructions, the voice wake-up method of the camera unit according to any one of claims 1 to 22 is realized.

28. A computer program product comprising a computer program, characterised in that, The computer program is executed by the processor to realize the voice wake-up method of the camera unit according to any one of claims 1 to 22.

Citation Information

Patent Citations

  • Voice wake-up method and device of camera unit, electronic equipment and storage medium

    CN121708923A

  • Large language model training method and interaction method oriented to multi-task dialogue

    CN116821290A

  • Multi-model cooperation method based on large-scale language model

    CN116976306A

  • Microscope control method and device driven by large language model, and electronic equipment

    CN117672222A

  • Smart home system voice control method, device and equipment and smart home system

    CN118248145A