Voice reply method, device, system, storage medium and product
By integrating a camera and dialogue model into the headset device, real-time image and voice information is acquired, and response voice information is generated and played, solving the problem of low efficiency in intelligent dialogue on terminal devices and achieving efficient information acquisition and improved user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2026-04-02
AI Technical Summary
In existing technologies, intelligent dialogue is mainly achieved through terminal devices, resulting in low dialogue efficiency.
A camera is deployed on the headset device to acquire target images in real time, and combined with the collected voice information to generate response voice information through a pre-trained dialogue model, which is then played directly on the headset device.
It improves the efficiency of intelligent dialogue, enabling users to quickly and easily obtain information about their surroundings, thus enhancing the user experience.
Smart Images

Figure CN2025088880_02042026_PF_FP_ABST
Abstract
Description
Voice reply method, device, system, storage medium and product
[0001] The present disclosure claims priority to the Chinese patent application No. 202411397888.7, filed on September 30, 2024, entitled "Voice reply method, device, system, storage medium and product", the entire content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] Embodiments of the present disclosure relate to the technical field of computer, and particularly relate to a voice reply method, device, system, storage medium and product. BACKGROUND
[0003] Intelligent dialogue is a mode of interaction between users and intelligent devices in the form of one question and one answer, and the intelligent dialogue is widely used in voice assistants, intelligent customer service and other scenarios.
[0004] At present, intelligent dialogue is mainly implemented by terminal devices, for example, a user inputs voice information or text information through the interface of a terminal device, and the terminal device outputs corresponding reply information, which has the problem of low dialogue efficiency. SUMMARY
[0005] Embodiments of the present disclosure provide a voice reply method, device, system, computer readable storage medium and product, which are used to solve the technical problem of low efficiency of intelligent dialogue implemented by terminal devices in the prior art.
[0006] In a first aspect, embodiments of the present disclosure provide a voice reply method applied to a headset device carrying a camera, and the method comprises:
[0007] collecting first inquiry voice information;
[0008] obtaining a target image collected by the camera;
[0009] obtaining first reply voice information of the first inquiry voice information, wherein the first reply voice information is information related to the target image, and the target image and the first inquiry voice information are processed by a pre-trained dialogue model to obtain the first reply voice information;
[0010] playing the first reply voice information.
[0011] In a second aspect, embodiments of the present disclosure provide a headset device, which comprises a headset device carrying a camera, and the headset device comprises:
[0012] a collecting unit configured to collect first inquiry voice information;
[0013] an obtaining unit configured to obtain a target image collected by the camera;
[0014] a processing unit configured to obtain first reply voice information of the first inquiry voice information, wherein the first reply voice information is information related to the target image, and the target image and the first inquiry voice information are processed by a pre-trained dialogue model to obtain the first reply voice information;
[0015] a playing unit configured to play the first reply voice information.
[0016] In a third aspect, an embodiment of the present application provides a voice reply system, comprising the earphone device and the terminal device of the second aspect, wherein:
[0017] the earphone device is configured to collect the first inquiry voice information, acquire the target image collected by the camera, and send the target image and the first inquiry voice information to the terminal device;
[0018] the terminal device is configured to process the target image and the first inquiry voice information by using a pre-trained dialogue model to obtain first reply voice information of the first inquiry voice information, and send the first reply voice information to the earphone device;
[0019] the earphone device is configured to receive the first reply voice information sent by the terminal device, and play the first reply voice information.
[0020] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory;
[0021] the memory stores computer execution instructions;
[0022] the processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the first aspect and various possible voice reply methods of the first aspect.
[0023] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores computer execution instructions, when the processor executes the computer execution instructions, the first aspect and various possible voice reply methods of the first aspect are realized.
[0024] In a sixth aspect, an embodiment of the present application provides a computer program product, comprising a computer program, when the computer program is executed by the processor, the first aspect and various possible voice reply methods of the first aspect are realized.
[0025] The voice reply method, device, system, storage medium and product provided by the embodiment include: collecting first inquiry voice information; obtaining a target image collected by a camera; obtaining first reply voice information of the first inquiry voice information, wherein the first reply voice information is information related to the target image, the target image and the first inquiry voice information are processed by a pre-trained dialogue model to obtain the first reply voice information; and playing the first reply voice information. First, the application realizes intelligent dialogue by using a headset, thereby improving the efficiency of intelligent dialogue. In addition, the application can quickly determine the first reply voice information by using a dialogue model. Furthermore, the target image is obtained to reply to the first inquiry voice information, so that the user's question about the real-time environment can be replied to. BRIEF DESCRIPTION OF DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art without creative labor.
[0027] FIG. 1 is an application scenario example diagram of a voice reply method provided by an embodiment of the present application;
[0028] FIG. 2 is a flowchart of a voice reply method provided by an embodiment of the present application;
[0029] FIG. 3 is a flowchart of another voice reply method provided by an embodiment of the present application;
[0030] FIG. 4 is a structural schematic diagram of a headset device provided by an embodiment of the present application;
[0031] FIG. 5 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present disclosure.
[0033] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the target user should be informed of the type, use range, use scenario, etc. of the personal information involved in the present disclosure and obtain the authorization of the target user in a proper manner according to relevant laws and regulations.
[0034] It can be understood that the above notification and target user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0035] In the related art, if a user needs to inquire about the related information of an object in the surrounding environment, the user needs to use a mobile terminal with a camera to take a picture of the object, obtain an image, and then use a related recognition application in the mobile terminal to scan the image, and display the related information of the object on the interface of the mobile terminal. If the user needs to know further information about the object, the user needs to search for the information in the application based on the related information, for example, when the user is in the current environment, the user sees a bunch of flowers and does not know the name of the flowers. The user uses a mobile terminal to take a picture of the flowers, obtains an image containing the flowers, and then the user uses a recognition application on the mobile terminal to scan the image and obtains the name of the flowers. Then the user needs to know the flowering period and planting place of the flowers, and the user needs to input questions about the flowers in the application to obtain the corresponding answers. It can be seen that the process of obtaining the related information of the object is complex and inefficient.
[0036] To solve the technical problem, the present disclosure provides a voice reply method, which deploys a camera on a headset device, controls the camera to obtain a target image in real time, and the headset device collects first inquiry voice information, obtains first reply voice information for the target image and the first inquiry voice information, and plays the first reply voice information through the headset device. This can enable the user to quickly and simply obtain the desired information and improve the conversation efficiency.
[0037] It should be noted that the voice reply method, device, system, storage medium and product provided by the present disclosure can be applied to any headset device configured with a camera.
[0038] The voice reply method provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0039] FIG. 1 is an example of an application scenario of a voice reply method provided in an embodiment of the present application. As shown in FIG. 1, a user wears a headset device, which includes a first headset and a second headset, one ear of the user wearing the first headset and the other ear wearing the second headset. The headset device carries a camera, which, in the embodiment of the present application, can be carried by the first headset and / or the second headset. In addition, the first headset and / or the second headset carries a microphone, which is used to collect inquiry voice information of the user. It can be understood that the camera carried by the headset device collects a target image, the microphone collects inquiry voice information of the user, and reply voice information can be generated according to the target image and the inquiry voice information of the user, which can be played through the headset device. The whole conversation process is simple and efficient, and the user experience is improved.
[0040] FIG. 1 is only an example of an application scenario, which can be used in any headset device configured with a camera, such as an in-ear headset, a semi-in-ear headset, a headset, a wired headset, a TWS (True Wireless Stereo) headset, etc. The type of the headset device is not limited in the present application.
[0041] FIG. 2 is a flowchart of a voice reply method provided in an embodiment of the present application, which can include the following steps:
[0042] S201, collect first inquiry voice information.
[0043] The first inquiry voice information is collected by the microphone of the headset device. It can be understood that one headset (the first headset or the second headset) of the headset device carries a microphone, which can collect the first inquiry voice information of the user in real time.
[0044] It can be understood that the collected first inquiry voice information is to be processed by a dialogue model by a preset application. In one embodiment, the preset application has been woken up, and the first inquiry voice information is any inquiry voice information input by the user, such as "what is in front of me?", "what is this flower called?", "what variety does this flower belong to?", or "what color does this flower have?", etc.
[0045] In another embodiment, the preset application has not been woken up, and the first inquiry voice information includes a preset wake-up word of the preset application and inquiry voice information, such as "XX, what is in front of me?", "XX, XX, what is in front of me?", "XX, take a look at what is in front of me".
[0046] The "XX" is the preset wake-up word of the preset application.
[0047] S202, obtain a target image collected by a camera.
[0048] The camera captures at least one target image or a series of continuous target images.
[0049] In an embodiment, before obtaining the target image captured by the camera, the method further comprises: waking up the camera in response to a wake-up operation on the camera; and obtaining the target image captured by the camera comprises: controlling the woken-up camera to capture the target image.
[0050] It can be understood that, in the case that the camera has been woken up, if the first inquiry voice information is captured, the camera is controlled to capture the target image after the first inquiry voice information is captured.
[0051] The camera can be woken up by a wake-up word or a touch control earphone device. The wake-up word of the camera can be the same as the wake-up word of a preset application. For example, the wake-up word is "XX", or "XX, look", "XX, scan", "XX, look again", "XX, take a closer look", or "XX, take a closer look".
[0052] In another embodiment, obtaining the target image captured by the camera comprises: determining a shooting intention of the first inquiry voice information by using a recognition model; and controlling the camera to capture the target image according to the shooting intention.
[0053] It can be understood that, after the first inquiry voice information is captured, the wake-up word of the camera in the first inquiry voice information is identified first, the camera is woken up, the shooting intention of the first inquiry voice information is analyzed by using the recognition model, for example, the first inquiry voice information is "What is the name of this flower", the shooting intention is to shoot the flower, and then the camera can be controlled to focus on the flower to capture the target image. In this way, an accurate target image can be obtained, which can be used for subsequent accurate reply to the user.
[0054] The recognition model is a network model deployed on the earphone device.
[0055] In the embodiments of the present application, the earphone device comprises a first earphone and a second earphone, the first earphone carries a first camera, the second earphone carries a second camera, and the target image comprises a target image captured by the first camera and a target image captured by the second camera.
[0056] It can be understood that, in an embodiment, at least one camera can be arranged on only one earphone, for example, only a first camera is arranged on a first earphone, and the first camera is required to capture the target image. In another embodiment, cameras can be arranged on both earphones, and the cameras of the two earphones capture the target image.
[0057] The target image collected by the first camera and the target image collected by the second camera are different angles of the same object, so that the determination of the first reply voice information can improve the reply accuracy of the first reply voice information.
[0058] S203, obtaining the first reply voice information of the first inquiry voice information.
[0059] The first reply voice information is information related to the target image, and the target image and the first inquiry voice information are processed by the pre-trained dialogue model to obtain the first reply voice information.
[0060] The dialogue model is an LLM (Large Language Model), and the dialogue model is a pre-trained multi-modal language model, which can process the target image and the natural language first inquiry voice information to obtain the first reply voice information. The first reply voice information is information related to the target image. For example, if the first inquiry voice information is "what is in front of me?", the target image includes "flowers", and the first reply voice information can be "there is a potted flower in front of me". If the first inquiry voice information is "what is this flower?", the target image includes "Chinese rose", and the first reply voice information can be "this is a Chinese rose".
[0061] In an embodiment, the dialogue model is deployed in the earphone device, and obtaining the first reply voice information of the first inquiry voice information includes: inputting the target image and the first inquiry voice information into the pre-trained dialogue model for processing to obtain the first reply voice information of the first inquiry voice information.
[0062] It can be understood that the dialogue model is deployed on the earphone device side, and then the earphone device inputs the first inquiry voice information and the target image into the dialogue model for processing to obtain the first reply voice information of the first inquiry voice information.
[0063] In the embodiments of the present application, the dialogue model can input the first inquiry voice information in voice format, or convert the first inquiry voice information into first inquiry text information through STT (Speech To Text) technology, and input the first inquiry text information into the dialogue model. The dialogue model can also output first reply text information, and the earphone device converts the first reply text information from text format to voice format by using TTS (Text To Speech) technology to obtain the first reply voice information.
[0064] In the embodiment of the present application, the earphone device can also be installed with a dialogue application, and a dialogue model can be deployed in a cloud server. The earphone device can upload the target image and the first inquiry voice information to the dialogue model of the cloud service through the dialogue application for processing to obtain the first reply voice information.
[0065] In another embodiment, obtaining the first reply voice information of the first inquiry voice information includes: sending the target image and the first inquiry voice information to a terminal device, the terminal device being configured to process the target image and the first inquiry voice information through a pre-trained dialogue model to obtain the first reply voice information of the first inquiry voice information; and receiving the first reply voice information sent by the terminal device.
[0066] It can be understood that when the dialogue model is deployed on the terminal device side, the earphone device sends the collected target image and the first inquiry voice information to the terminal device, and the terminal device processes the target image and the first inquiry voice information through the dialogue model to obtain the first reply voice information.
[0067] In the embodiment of the present application, the terminal device can also be installed with a dialogue application, and a dialogue model can be deployed in a cloud server. The terminal device can upload the target image and the first inquiry voice information to the dialogue model of the cloud service through the dialogue application for processing to obtain the first reply voice information, and then send the first reply voice information to the earphone device for playing.
[0068] In summary, the present application can adopt various ways to process the target image and the first inquiry voice information through the dialogue model to obtain the first reply voice information.
[0069] S204, playing the first reply voice information.
[0070] In the embodiment of the present application, the first reply voice information can be played on one earphone (the first earphone or the second earphone), or the first reply voice information can be played on both the first earphone and the second earphone.
[0071] In summary, the present application deploys a camera on the earphone device, controls the camera to acquire the target image in real time, and collects the first inquiry voice information by the earphone device. The first reply voice information is obtained for the target image and the first inquiry voice information, and the first reply voice information is played through the earphone device. This can enable the user to quickly and simply obtain the desired information and improve the dialogue efficiency.
[0072] FIG. 3 is a flowchart of another voice reply method provided by an embodiment of the present application. The voice reply method can include the following steps:
[0073] S301, collecting first inquiry voice information.
[0074] The specific implementation process of this step refers to S201, which will not be repeated here.
[0075] S302, acquiring the target image collected by the camera.
[0076] The specific implementation process of this step refers to S202, which will not be repeated here.
[0077] S303, sending the target image and the first inquiry voice information to the terminal device.
[0078] In the embodiments of the present application, the earphone device and the terminal device are in communication connection, wherein the earphone device sends the target image and the first inquiry voice information to the terminal device through Bluetooth or WIFI (mobile hotspot). When the earphone device sends the target image and the first inquiry voice information to the terminal device through WIFI, the earphone device can connect the hotspot of the terminal device.
[0079] Among them, the user can set the camera of the single earphone (the first earphone or the second earphone) to capture the target image, and the user can also set the cameras of the double earphones (the first earphone and the second earphone) to capture the target image.
[0080] In one embodiment, sending the target image to the terminal device comprises: sending the target image captured by the first camera to the terminal device by using the first earphone; and sending the target image captured by the second camera to the terminal device by using the second earphone. When the cameras of the double earphones are set to capture the target image, one earphone (for example, the first earphone) can be set as the master earphone, and the other earphone (for example, the second earphone) can be set as the slave earphone.
[0081] It can be understood that the first earphone and the terminal device are in communication connection, and the second earphone and the terminal device are in communication connection. Therefore, the first earphone and the second earphone can send the target image to the terminal device separately.
[0082] Among them, if the first earphone side carries a microphone, the first earphone sends the target image captured by the first camera and the first inquiry voice information collected by the first earphone to the terminal device, and the second earphone sends the target image captured by the second camera to the terminal device. Among them, if the second earphone side carries a microphone, the first earphone sends the target image captured by the first camera to the terminal device, and the second earphone sends the target image captured by the second camera and the second inquiry voice information collected by the second earphone to the terminal device.
[0083] In the embodiments of the present application, the first earphone and the second earphone are in communication with the terminal device separately, without considering the connection between the earphones, wherein the earphone carrying the microphone is the master earphone.
[0084] In another embodiment, the method for sending the target image to the terminal device comprises: the first earphone receiving the target image captured by the second camera of the second earphone; and the first earphone sending the target image captured by the first camera and the target image captured by the second camera to the terminal device.
[0085] It can be understood that the first earphone is communicatively connected with the terminal device, the second earphone is not communicatively connected with the terminal device, and the first earphone is communicatively connected with the second earphone. In this embodiment, the first earphone is the master earphone, and the second earphone is the slave earphone. In this case, the second earphone sends the target image captured by the second camera to the first earphone, and the first earphone sends the target image captured by the first camera and the target image captured by the second camera to the terminal device.
[0086] Further, in the case where the first earphone is communicatively connected with the terminal device, the second earphone is not communicatively connected with the terminal device, and the first earphone is communicatively connected with the second earphone, if the first earphone is provided with a microphone, the first earphone collects the first inquiry voice information, and the first earphone sends the first inquiry voice information to the terminal device. If the second earphone is provided with a microphone, the second earphone collects the first inquiry voice information, the second earphone sends the first inquiry voice information to the first earphone, and the first earphone sends the first inquiry voice information to the terminal device.
[0087] In the embodiment of the present application, after the first earphone receives the target image sent by the second earphone, the first earphone sends the target image to the terminal device. Only one earphone needs to be communicatively connected with the terminal device. The link for transmitting data from one earphone to the terminal device is short.
[0088] Further, in the present application, the target image captured by the camera of the master earphone can be sent to the terminal device first, and then the target image captured by the camera of the slave earphone can be sent to the terminal device, or the target images can be sent simultaneously, which is not limited herein.
[0089] In addition, the first inquiry voice information is encoded before being sent to the terminal device, for example, the first inquiry voice information is encoded by using OPUS encoding (a kind of sound encoding mode). After the terminal device receives the encoded first inquiry voice information, the first inquiry voice information is decoded by using OPUS decoding (a kind of sound decoding mode) and then input into the question and answer model for processing.
[0090] S304, receiving the first reply voice information sent by the terminal device.
[0091] In one embodiment, the first earphone is used to receive the first reply voice information sent by the terminal device, and the first earphone forwards the first reply voice information to the second earphone.
[0092] In the case that the first earphone and the terminal device are communicatively connected and the first earphone and the second earphone are communicatively connected, the terminal device sends the first reply voice information to the first earphone, and the first earphone forwards the first reply voice information to the second earphone.
[0093] In another embodiment, the first earphone receives the first reply voice information sent by the terminal device, and the second earphone receives the first reply voice information sent by the terminal device.
[0094] In the case that the first earphone and the terminal device are communicatively connected and the first earphone and the second earphone are communicatively connected, the terminal device sends the first reply voice information to the first earphone, and the first earphone forwards the first reply voice information to the second earphone.
[0095] S305, playing the first reply voice information.
[0096] In the embodiment of the present application, the terminal device can encode the first reply voice information and send it to the earphone device. After the earphone device receives the encoded first reply voice information, it decodes and plays the first reply voice information.
[0097] In the embodiment of the present application, the terminal device can encode the first reply voice information and send it to the earphone device. After the earphone device receives the encoded first reply voice information, it decodes and plays the first reply voice information.
[0098] S306, collecting the second inquiry voice information.
[0099] It can be understood that the user can continue to inquire in response to the first reply voice information played by the earphone device, realizing multi-round dialogue.
[0100] S307, obtaining the second reply voice information of the second inquiry voice information.
[0101] In the embodiment of the present application, the terminal device can encode the first reply voice information and send it to the earphone device. After the earphone device receives the encoded first reply voice information, it decodes and plays the first reply voice information.
[0102] In the embodiment of the present application, the user can realize polling dialogue in response to the target image. For example, if the first inquiry voice information is "what is in front of me?", the target image includes "flowers", and the first reply voice information can be "there are potted flowers in front of me", and the second inquiry voice information is "what is the name of this flower", then the second inquiry voice information and the target image are input into the dialogue model for processing to obtain the second reply voice information.
[0103] In the embodiment of the present application, the terminal device can encode the first reply voice information and send it to the earphone device. After the earphone device receives the encoded first reply voice information, it decodes and plays the first reply voice information.
[0104] S308, playing the second reply voice information.
[0105] In an embodiment, after the multi-round dialogue for the target image ends, the earphone device can further collect other inquiry voice information and other images to start a new round of dialogue.
[0106] In the embodiments of the present application, through the communication connection between the earphone device and the terminal device, the earphone device can be endowed with the function of intelligent dialogue, the efficiency of implementing intelligent dialogue is improved, and the user experience is improved.
[0107] FIG. 4 is a structural schematic diagram of an earphone device provided by an embodiment of the present application. The earphone device carries a camera. The earphone device 40 can include the following units:
[0108] The acquisition unit 41 is configured to acquire first inquiry voice information.
[0109] The acquisition unit 42 is configured to acquire a target image collected by the camera.
[0110] The processing unit 43 is configured to obtain first reply voice information of the first inquiry voice information, wherein the first reply voice information is information related to the target image, and the target image and the first inquiry voice information are processed by a pre-trained dialogue model to obtain the first reply voice information.
[0111] The playing unit 44 is configured to play the first reply voice information.
[0112] In some embodiments, the earphone device further includes a waking unit (not shown) configured to wake up the camera in response to a waking operation on the camera before the target image collected by the camera is acquired.
[0113] The acquisition unit 42 is specifically configured to control the woken-up camera to collect the target image.
[0114] In some embodiments, the acquisition unit 42 is specifically configured to determine a shooting intention of the first inquiry voice information through a recognition model, and control the camera to collect the target image according to the shooting intention.
[0115] In some embodiments, the dialogue model is deployed in the earphone device, and the processing unit 43 is specifically configured to input the target image and the first inquiry voice information into the pre-trained dialogue model for processing to obtain the first reply voice information of the first inquiry voice information.
[0116] In some embodiments, the processing unit 43 is specifically configured to send the target image and the first inquiry voice information to a terminal device, the terminal device is configured to process the target image and the first inquiry voice information through a pre-trained dialogue model to obtain the first reply voice information of the first inquiry voice information, and receive the first reply voice information sent by the terminal device.
[0117] In some embodiments, the earphone device comprises: a first earphone and a second earphone, the first earphone carries a first camera, and the second earphone carries a second camera; and the target image comprises: a target image captured by the first camera and a target image captured by the second camera.
[0118] In some embodiments, when the processing unit 43 sends the target image to the terminal device, specifically: the first earphone sends the target image captured by the first camera to the terminal device; and the second earphone sends the target image captured by the second camera to the terminal device.
[0119] In some embodiments, when the processing unit 43 sends the target image to the terminal device, specifically: the first earphone receives the target image captured by the second camera sent by the second earphone; and the first earphone sends the target image captured by the first camera and the target image captured by the second camera to the terminal device.
[0120] In some embodiments, when the processing unit 43 receives the first reply voice information sent by the terminal device, specifically: the first earphone receives the first reply voice information sent by the terminal device; and the first earphone forwards the first reply voice information to the second earphone.
[0121] In some embodiments, when the processing unit 43 receives the first reply voice information sent by the terminal device, specifically: the first earphone receives the first reply voice information sent by the terminal device; and the second earphone receives the first reply voice information sent by the terminal device.
[0122] In some embodiments, after playing the first reply voice information,
[0123] The acquisition unit 41 is further configured to acquire second inquiry voice information.
[0124] The processing unit 43 is further configured to obtain second reply voice information of the second inquiry voice information, wherein the second reply voice information is information related to the target image, and the target image and the second inquiry voice information are processed by the dialogue model to obtain the second reply voice information.
[0125] The playing unit 44 is further configured to play the second reply voice information.
[0126] In order to realize the above-mentioned embodiments, the embodiments of the present disclosure further provide a voice reply system comprising the earphone device and the terminal device mentioned above.
[0127] Wherein:
[0128] The earphone device is configured to acquire first inquiry voice information; acquire a target image captured by a camera; and send the target image and the first inquiry voice information to a terminal device.
[0129] The terminal device is configured to process the target image and the first inquiry voice information through a pre-trained dialogue model to obtain first reply voice information of the first inquiry voice information, and send the first reply voice information to the earphone device.
[0130] The earphone device is configured to receive the first reply voice information sent by the terminal device and play the first reply voice information.
[0131] To implement the above-mentioned embodiments, the embodiments of the present disclosure further provide a computer readable storage medium, which stores computer execution instructions. When a processor executes the computer execution instructions, the voice reply method of any one of the above-mentioned embodiments is implemented.
[0132] To implement the above-mentioned embodiments, the embodiments of the present disclosure further provide a computer program product, which includes a computer program. When the computer program is executed by a processor, the voice reply method of any one of the above-mentioned embodiments is implemented.
[0133] To implement the above-mentioned embodiments, the embodiments of the present disclosure further provide an electronic device, which includes a processor and a memory.
[0134] The memory stores computer execution instructions.
[0135] The processor executes the computer execution instructions stored in the memory, so that the processor executes the voice reply method of any one of the above-mentioned embodiments.
[0136] FIG. 5 is a structural schematic diagram of an electronic device provided by the embodiments of the present disclosure. The electronic device 50 can be a terminal device or a server. The terminal device can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a personal digital assistant (PDA), a tablet computer (PAD), a portable media player (PMP), a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 5 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.
[0137] As shown in FIG. 5, the electronic device 50 can include a processing device (e.g., a central processor, a graphics processor, etc.) 51 that can perform various appropriate actions and processes according to programs stored in a Read Only Memory (ROM) 52 or loaded into a Random Access Memory (RAM) 53 from a storage device 58. Various programs and data required for the operation of the electronic device 50 are also stored in the RAM 53. The processing device 51, the ROM 52, and the RAM 53 are connected to each other through a bus 54. An Input / Output (I / O) interface 55 is also connected to the bus 54.
[0138] Generally, the following devices can be connected to the I / O interface 55: input devices 56 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 57 including, for example, a Liquid Crystal Display (LCD), a speaker, a vibrator, etc.; storage devices 58 including, for example, a tape, a hard disk, etc.; and communication devices 59. The communication devices 59 can allow the electronic device 50 to communicate wirelessly or wired with other devices to exchange data. Although FIG. 5 shows the electronic device 50 with various devices, it should be understood that all of the shown devices are not required to be implemented or possessed. More or less devices can be alternatively implemented or possessed.
[0139] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 59, or installed from the storage devices 58, or installed from the ROM 52. When the computer program is executed by the processing device 51, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0140] It should be noted that the computer-readable medium in the above disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.
[0141] The computer-readable medium described above can be contained in the electronic device described above; or can exist separately and not be assembled into the electronic device.
[0142] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.
[0143] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0144] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0145] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.
[0146] The functions described in this specification can be performed at least in part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0147] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0148] In a first aspect, according to one or more embodiments of the present disclosure, a voice reply method is provided, applied to a headset device carrying a camera, the method comprising: collecting first inquiry voice information; obtaining a target image collected by the camera; obtaining first reply voice information of the first inquiry voice information, wherein the first reply voice information is information related to the target image, and the target image and the first inquiry voice information are processed by a pre-trained dialogue model to obtain the first reply voice information; and playing the first reply voice information.
[0149] According to one or more embodiments of the present disclosure, before obtaining the target image collected by the camera, the method further comprises: in response to a wake-up operation on the camera, waking up the camera; and obtaining the target image collected by the camera, comprising: controlling the woken-up camera to collect the target image.
[0150] According to one or more embodiments of the present disclosure, obtaining the target image collected by the camera comprises: determining a shooting intention of the first inquiry voice information through a recognition model; and controlling the camera to collect the target image according to the shooting intention.
[0151] According to one or more embodiments of the present disclosure, the dialogue model is deployed in the headset device, and obtaining the first reply voice information of the first inquiry voice information comprises: inputting the target image and the first inquiry voice information into the pre-trained dialogue model for processing to obtain the first reply voice information of the first inquiry voice information.
[0152] According to one or more embodiments of the present disclosure, obtaining the first reply voice information of the first inquiry voice information comprises: sending the target image and the first inquiry voice information to a terminal device, the terminal device being configured to process the target image and the first inquiry voice information through a pre-trained dialogue model to obtain the first reply voice information of the first inquiry voice information; and receiving the first reply voice information sent by the terminal device.
[0153] According to one or more embodiments of the present disclosure, the earphone device comprises: a first earphone and a second earphone, the first earphone carries a first camera, and the second earphone carries a second camera; and the target image comprises: a target image captured by the first camera and a target image captured by the second camera.
[0154] According to one or more embodiments of the present disclosure, the method for sending the target image to the terminal device comprises: sending, by the first earphone, the target image captured by the first camera to the terminal device; and sending, by the second earphone, the target image captured by the second camera to the terminal device.
[0155] According to one or more embodiments of the present disclosure, the method for sending the target image to the terminal device comprises: receiving, by the first earphone, the target image captured by the second camera sent by the second earphone; and sending, by the first earphone, the target image captured by the first camera and the target image captured by the second camera to the terminal device.
[0156] According to one or more embodiments of the present disclosure, the method for receiving the first reply voice information sent by the terminal device comprises: receiving, by the first earphone, the first reply voice information sent by the terminal device; and forwarding, by the first earphone, the first reply voice information to the second earphone.
[0157] According to one or more embodiments of the present disclosure, the method for receiving the first reply voice information sent by the terminal device comprises: receiving, by the first earphone, the first reply voice information sent by the terminal device; and receiving, by the second earphone, the first reply voice information sent by the terminal device.
[0158] According to one or more embodiments of the present disclosure, after the first reply voice information is played, the method further comprises: capturing second inquiry voice information; obtaining second reply voice information of the second inquiry voice information, wherein the second reply voice information is information related to the target image, and the target image and the second inquiry voice information are processed by a dialogue model to obtain the second reply voice information; and playing the second reply voice information.
[0159] In a second aspect, according to one or more embodiments of the present disclosure, an earphone device is provided, the earphone device carries a camera, and the earphone device comprises:
[0160] The capturing unit is configured to capture first inquiry voice information.
[0161] The obtaining unit is configured to obtain a target image captured by the camera.
[0162] The processing unit is configured to obtain first reply voice information of the first inquiry voice information, wherein the first reply voice information is information related to the target image, and the target image and the first inquiry voice information are processed by a dialogue model trained in advance to obtain the first reply voice information.
[0163] a playing unit, configured to play the first reply voice information.
[0164] In a third aspect, an earphone device and a terminal device are provided according to one or more embodiments of the present disclosure, and the earphone device and the terminal device are used in the voice reply system according to the second aspect.
[0165] The earphone device is configured to collect first inquiry voice information, acquire a target image collected by a camera, and send the target image and the first inquiry voice information to the terminal device.
[0166] The terminal device is configured to process the target image and the first inquiry voice information by using a pre-trained dialogue model to obtain first reply voice information of the first inquiry voice information, and send the first reply voice information to the earphone device.
[0167] The earphone device is configured to receive the first reply voice information sent by the terminal device, and play the first reply voice information.
[0168] In a fourth aspect, an electronic device is provided according to one or more embodiments of the present disclosure, and the electronic device includes at least one processor and a memory.
[0169] The memory stores computer-executable instructions.
[0170] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the voice reply method according to the first aspect and various possible designs of the first aspect.
[0171] In a fifth aspect, a computer-readable storage medium is provided according to one or more embodiments of the present disclosure, and the computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the voice reply method according to the first aspect and various possible designs of the first aspect is implemented.
[0172] In a sixth aspect, a computer program product is provided according to one or more embodiments of the present disclosure, and the computer program product includes a computer program, and when a processor executes the computer program, the voice reply method according to the first aspect and various possible designs of the first aspect is implemented.
[0173] The above description is merely preferred embodiments of the present disclosure and a description of principles of applied technologies. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the disclosed concept. For example, the above technical features are replaced with technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0174] Further, although operations are depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order, and that certain features of the disclosure can be performed in parallel or concurrently with one another. Similarly, while operations have been depicted as following a sequential flow, in other implementations, the operations can be performed in different orders or concurrently. Additionally, the inclusion of certain features does not mean that those features are required. Furthermore, to the extent that the terms "includes", "containing", "has", "having", and / or variants thereof are used in either the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "comprising" as an open transition term without precluding any additional or
[0175] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. An earphone device, characterized by, The earphone device comprises a first earphone and a second earphone, the first earphone carries a camera and / or the second earphone carries a camera.
2. The earphone device of claim 1, wherein, The earphone device comprises one of an in-ear earphone, a semi-in-ear earphone, a headphone, a wired earphone, and a truly wireless earphone.
3. The earphone device of claim 1, wherein, The earphone device is used for collecting a target image, the first earphone carries a first camera, and the target image comprises a target image collected by the first camera and a target image collected by the second camera.
4. The earphone device of claim 3, wherein, The first earphone is used for sending the target image collected by the first camera to a terminal device, and the second earphone is used for sending the target image collected by the second camera to the terminal device.
5. The earphone device of claim 3, wherein, The first earphone is used for receiving the target image collected by the second camera sent by the second earphone, and sending the target image collected by the first camera and the target image collected by the second camera to the terminal device.
6. The earphone device of claim 1, wherein, The earphone device is used for collecting a target image, the first earphone carries a first camera, and the target image comprises a target image collected by the first camera; or, the second earphone carries a second camera, and the target image comprises a target image collected by the second camera.
7. The earphone device according to claim 3 or 6, characterized in that, The earphone device further comprises a microphone capable of collecting first inquiry voice information, and the target image and the first inquiry voice information are input to a pre-trained dialogue model by the earphone for processing.
8. The earphone device of claim 7, wherein, The earphone device is deployed with the dialogue model, or the earphone device is installed with a dialogue application, the cloud server is deployed with the dialogue model, and the earphone device uploads the target image and the first inquiry voice information to the dialogue model of the cloud server for processing through the dialogue application.
9. The earphone device of claim 7, wherein, The earphone device is in communication connection with a terminal device, the terminal device is deployed with the dialogue model, or the terminal device is installed with a dialogue application, the cloud server is deployed with the dialogue model, and the terminal device uploads the target image and the first inquiry voice information to the dialogue model of the cloud server for processing through the dialogue application.
10. A voice reply method, characterized by, The method is applied to an earphone device carrying a camera, and the method comprises the following steps: collecting first inquiry voice information; acquiring a target image collected by the camera; obtaining first reply voice information of the first inquiry voice information, wherein the first reply voice information is information related to the target image, and the target image and the first inquiry voice information are processed by a pre-trained dialogue model to obtain the first reply voice information; playing the first reply voice information.
11. The voice reply method of claim 10, wherein, Before the step of acquiring the target image collected by the camera, the method further comprises the following steps: in response to a wake-up operation on the camera, waking up the camera; the step of acquiring the target image collected by the camera comprises the following step:
12. The voice reply method of claim 10, wherein, controlling the woken-up camera to collect the target image. The step of acquiring the target image collected by the camera comprises the following steps: determining a shooting intention of the first inquiry voice information through a recognition model; controlling the camera to collect the target image according to the shooting intention.
13. The voice reply method of claim 10, wherein, The dialogue model is deployed in the earphone device, and obtaining the first reply voice information of the first inquiry voice information includes: The target image and the first inquiry voice information are input into a pre-trained dialogue model for processing to obtain the first reply voice information of the first inquiry voice information.
14. The voice reply method of claim 10, wherein, The first reply voice information of the first inquiry voice information includes: The target image and the first inquiry voice information are sent to a terminal device, and the terminal device is configured to process the target image and the first inquiry voice information through a pre-trained dialogue model to obtain the first reply voice information of the first inquiry voice information. The first reply voice information sent by the terminal device is received.
15. The voice reply method according to any one of claims 10 to 14, characterized by, The earphone device includes a first earphone and a second earphone, the first earphone carries a first camera, the second earphone carries a second camera, and the target image includes a target image captured by the first camera and a target image captured by the second camera.
16. The voice reply method of claim 15, wherein, Sending the target image to a terminal device includes: The first earphone is used to send the target image captured by the first camera to the terminal device; The second earphone is used to send the target image captured by the second camera to the terminal device.
17. The voice reply method of claim 15, wherein, Sending the target image to a terminal device includes: The first earphone receives the target image captured by the second camera sent by the second earphone; The first earphone is used to send the target image captured by the first camera and the target image captured by the second camera to the terminal device.
18. The voice reply method of claim 15, wherein, The first reply voice information sent by the terminal device is received. The first earphone is used to receive the first reply voice information sent by the terminal device; The first earphone forwards the first reply voice information to the second earphone.
19. The voice reply method of claim 15, wherein, The first reply voice information sent by the terminal device is received. The first earphone is used to receive the first reply voice information sent by the terminal device; The second earphone is used to receive the first reply voice information sent by the terminal device.
20. The voice reply method of any one of claims 10 to 14, wherein, After playing the first reply voice information, the method further includes: Capturing second inquiry voice information; Obtaining second reply voice information of the second inquiry voice information, wherein the second reply voice information is information related to the target image, and the target image and the second inquiry voice information are processed by the dialogue model to obtain the second reply voice information; Playing the second reply voice information.
21. An earphone device, comprising: The earphone device carries a camera, and the earphone device includes: A capturing unit configured to capture first inquiry voice information; An obtaining unit configured to obtain a target image captured by the camera; A processing unit configured to obtain first reply voice information of the first inquiry voice information, wherein the first reply voice information is information related to the target image, and the target image and the first inquiry voice information are processed by a pre-trained dialogue model to obtain the first reply voice information; A playing unit configured to play the first reply voice information.
22. A voice reply system characterized by The earphone device and the terminal device of claim 21 are included. The earphone device is configured to collect first inquiry voice information, acquire a target image collected by the camera, and send the target image and the first inquiry voice information to the terminal device. The terminal device is configured to process the target image and the first inquiry voice information by using a pre-trained dialogue model to obtain first reply voice information of the first inquiry voice information, and send the first reply voice information to the earphone device. The earphone device is configured to receive the first reply voice information sent by the terminal device, and play the first reply voice information.
23. An electronic device, comprising: Comprise: a processor and a memory; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory, so that the processor executes the voice reply method according to any one of claims 10 to 20.
24. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and when the processor executes the computer execution instructions, the voice reply method according to any one of claims 10 to 20 is realized.
25. A computer program product comprising a computer program, characterised in that, The computer program is executed by the processor to realize the voice reply method according to any one of claims 10 to 20.
Citation Information
Patent Citations
Voice reply method, device and system, storage medium and product
CN121764437A
Intelligent earphone system and control method thereof
CN104244132A
Information identification method, earphone and terminal equipment
CN110248269A
Intelligent earphone device with computer vision
CN112073866A
Interaction method and device based on earphone camera module, equipment and storage medium
CN119299831A