Head-mounted device

By combining a head-mounted device with a camera and a processor, and using a convolutional neural network to identify environmental image feature information and output voice, the problem of existing voice question-and-answer systems being unable to recognize image features is solved, and the environmental cognition ability and Internet experience of blind users are improved.

CN115357748BActive Publication Date: 2025-10-21TENCENT TECH SHANGHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210977146.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-01-17
Publication Date
2025-10-21
Estimated Expiration
2037-01-17

AI Technical Summary

Technical Problem

Existing voice question-answering systems are unable to recognize image feature information, resulting in blind users being unable to understand non-textual information in the environment, limiting their environmental cognition ability.

Method used

A head-mounted device, combined with a camera and a processor, receives query information through a voice interaction module, obtains feature information of the environmental image, and uses a query model trained by a convolutional neural network to identify target feature information and play voice query results.

Benefits of technology

It realizes the recognition and voice output of environmental image feature information, improves the intelligence of the voice query system, helps blind users understand the surrounding environment and web page information, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115357748B_ABST
    Figure CN115357748B_ABST
Patent Text Reader

Abstract

The application discloses a kind of head-mounted devices.Therein, the device includes: voice interaction module, for receiving the voice query information of current identification object, wherein, current identification object wears head-mounted device, and voice query information carries the query keyword for inquiring the target environment where current identification object is located;Camera, for shooting the environmental image observed by current identification object under current observation visual angle;Processor, for obtaining the feature information of environmental image, and inquiring the target feature information matched with query keyword from feature information, wherein, feature information is used to indicate the object in environmental image;Loudspeaker, for playing voice query result in the case of inquiring target feature information, wherein, voice query result is used to indicate target object in environmental image, and target object is indicated by target feature information.The application solves the technical problem that voice query system is not intelligent due to the inability to identify image feature information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent recognition, and in particular to a head-mounted device. Background Art

[0002] With the advancement of technology, the development of intelligent technology tends to meet the life and work needs of more and more people. For example, some voice question-answering systems can meet the information query needs of the blind. Currently, some blind question-answering systems on the market only solve knowledge-level problems. Generally, through voice interaction, the question is first converted into text through Speech2Text, and then a knowledge base search is performed to return the answer to the corresponding question. Finally, the answer is told to the questioner through the Text2Speech process.

[0003] The biggest problem with some existing products is that they only address text semantics, recognizing only text or language. However, no product currently addresses the environmental cognition challenges faced by blind people. The world is rich and diverse, and blind people need to understand it, too. For example, blind people currently use screen readers to browse the internet, but these apps can only interpret text; they cannot read images interspersed within text.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0005] An embodiment of the present invention provides a head-mounted device to at least solve the technical problem that a voice query system is not intelligent due to the inability to recognize image feature information.

[0006] According to one aspect of an embodiment of the present invention, a head-mounted device is provided, including: a voice interaction module for receiving voice query information of a currently identified object, wherein the currently identified object wears the head-mounted device, and the voice query information carries query keywords for querying a target environment where the currently identified object is located; a camera for capturing an environmental image of the currently identified object under a current observation angle; a processor for acquiring feature information of the environmental image, and querying target feature information matching the query keywords from the feature information, wherein the feature information is used to represent an object in the environmental image; and a speaker for playing a voice query result when the target feature information is found, wherein the voice query result is used to indicate the target object in the environmental image, and the target object is represented by the target feature information.

[0007] In an embodiment of the present invention, a head-mounted device is used to receive voice query information and obtain feature information of an image to be identified. Then, target feature information matching the query keyword is searched from the feature information. When the target feature information is found, the voice query result is played, thereby achieving the purpose of outputting the query result by voice after identifying the feature information in the image, thereby achieving the technical effect of improving the intelligence of the voice query system, and further solving the technical problem of the voice query system being unintelligent due to the inability to identify image feature information. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0009] Figure 1 is a schematic diagram of a hardware environment of a voice query method according to an embodiment of the present invention;

[0010] Figure 2 is a flow chart of an optional voice query method according to an embodiment of the present invention;

[0011] Figure 3 2 is a schematic diagram of the principle of a voice query system according to an embodiment of the present invention;

[0012] Figure 4 is a schematic diagram of an optional voice query device according to an embodiment of the present invention;

[0013] Figure 5 is a schematic diagram of an optional voice query device according to an embodiment of the present invention;

[0014] Figure 6 is a structural block diagram of a terminal according to an embodiment of the present invention;

[0015] Figure 7 is a schematic diagram of an optional head-mounted device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0016] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0017] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0018] Example 1

[0019] According to an embodiment of the present invention, a method embodiment of a voice query is provided.

[0020] Optionally, in this embodiment, the above voice query method can be applied to Figure 1 In the hardware environment shown in FIG. 1 , which is composed of a server 102 and a terminal 104. Figure 1 As shown, server 102 is connected to terminal 104 via a network, which includes but is not limited to a wide area network, a metropolitan area network, or a local area network. Terminal 104 is not limited to a PC, a mobile phone, a tablet computer, etc. The voice query method of the embodiment of the present invention can be executed by server 102, by terminal 104, or by both server 102 and terminal 104. The voice query method of the embodiment of the present invention can also be executed by a client installed on terminal 104.

[0021] In one application scenario, a user, especially a blind user, walking on the road and unable to understand their surroundings can use voice to query their surroundings. For example, if a user asks, "What's on the ground?" Upon receiving the user's voice query, the system captures an image of the surrounding environment, extracts feature information from the image, and detects that there are pigeons and bicycles on the ground. This information is then announced to the user via voice. This allows the blind user to keep abreast of their surroundings, improving the device's intelligence and user experience.

[0022] Figure 2 is a flow chart of an optional voice query method according to an embodiment of the present invention, such as Figure 2 As shown, the method may include the following steps:

[0023] Step S202: receiving voice query information, wherein the voice query information is used to indicate a query keyword.

[0024] Step S204: Acquire feature information of the image to be identified, wherein the feature information is used to represent the object in the image to be identified.

[0025] Step S206: Query the feature information for target feature information that matches the query keyword.

[0026] Step S208 : When the target feature information is found, the voice query result is played, wherein the voice query result is used to indicate the target object in the image, and the target object is represented by the target feature information.

[0027] Through the above steps S202 to S208, by receiving voice query information and obtaining feature information of the image to be identified, and then searching for target feature information matching the query keyword from the feature information, when the target feature information is found, the voice query result is played. Since the feature information of the image can be identified based on the voice query information and announced in voice form, the technical problem of the voice query system being unintelligent due to the inability to identify image feature information is solved, thereby achieving the technical effect of outputting the query result by voice after identifying the feature information in the image.

[0028] In the technical solution provided in step S202, the voice query method of the embodiment of the present invention can be implemented by a voice query system. Optional application scenarios include: a user wearing or carrying a device integrated with the voice query system to enable voice questions and answers about their surroundings; or a user browsing webpage information through a terminal equipped with the voice query system to enable voice questions and answers about images on the webpage. The received voice query information can be a voice query information issued by the user, for example, the user asks the voice query system: "What's on the ground?" or "What's on the webpage?" The voice query information contains query keywords, such as "on the ground," "webpage," and "what." The voice query information may also contain other adverbs, modal particles, etc.

[0029] In the technical solution provided in step S204, the image to be identified can be a real-time image of the surrounding environment or an image on a webpage. The feature information is used to represent the object in the image to be identified. The object in the image can be an object presented in the image, such as a table, a chair, a bird, the sky, etc., or other features of the object, such as color, size, etc. The step of obtaining the feature information of the image to be identified can be after receiving the voice query information, or before receiving the voice query information. The image to be identified can be obtained in real time or at predetermined intervals, such as every second, after the image to be identified is obtained.

[0030] In the technical solution provided in step S206, after obtaining the feature information of the image to be identified, the feature information is searched for target feature information that matches the query keyword. An image to be identified may contain multiple feature information. Since not every feature information is of interest to the user, the feature information of the image to be identified is searched for target feature information that matches the query keyword. During the query process, the feature information can be matched with the keyword and searched in a variety of ways.

[0031] In the technical solution provided in step S208, if the target feature information is found, the voice query result is played. The voice query result is used to indicate the target object in the image, which is represented by the target feature information. If the target feature information is found, the voice query result is played. For example, after receiving the user's question "What's on the ground?", if the image to be identified shows that there are pigeons on the ground, it is determined that the target feature information has been found, and the voice query result is output, that is, the word "pigeon" can be played by voice.

[0032] As an optional embodiment, querying target feature information that matches the query keyword from the feature information includes: inputting the feature information and the query keyword into a trained query model, wherein the trained query model is used to query target feature information that matches the query keyword from the feature information; when the query model outputs the target feature information, determining that the target feature information has been queried.

[0033] Querying target feature information that matches the query keyword from the feature information can be achieved through a query model. The feature information and the query keyword are input into a query model that has been trained. The query model can be pre-trained with a neural network. If the feature information and the query keyword are input into the query model that has been trained, and the query model outputs the target feature information, it indicates that the target feature information has been queried. At this time, it can be determined that the target feature information has been queried. If the feature information and the query keyword are input into the query model that has been trained, and the query model does not output the target feature information, or reports an error, it indicates that the target feature information has not been queried. At this time, it is determined that the target feature information has not been queried.

[0034] As an optional embodiment, before querying the target feature information that matches the query keyword from the feature information, a pre-set query model is trained through a convolutional neural network to obtain a trained query model, wherein, during the training process, the object features in the pre-obtained multiple images and the pre-obtained information features are used as inputs of the query model, the object features are used to represent the objects in the multiple images, and the information features are used to represent the query questions in a predetermined query question set.

[0035] When training a pre-set query model, a convolutional neural network can be used. For example, a residual computer network (ResNet) convolutional neural network can be used to train the pre-set query model. During model training, the query model inputs are: pre-obtained object features from multiple images and pre-obtained information features. The object features can represent objects in the multiple images, the multiple images can be images from the ImageNet dataset, and the information features can be features of questions from a predetermined set of query questions.

[0036] As an optional embodiment, a pre-set query model is trained through a convolutional neural network to obtain a trained query model, which can be: obtaining the correlation between object features and information features; continuously adjusting the values ​​of parameters in the query model until the highest correlation is obtained, wherein the values ​​of the parameters in the trained query model are the values ​​of the parameters when the correlation is the highest.

[0037] The pre-set query model is trained by a convolutional neural network. The trained query model can be obtained through the following steps: first, the correlation between the object features and the information features is obtained; while adjusting the parameter values ​​in the query model, the correlation between the object features and the information features is continued to be obtained to obtain the value with the maximum correlation between the object features and the information features; when the correlation value is the maximum, the values ​​of these parameters are used as the parameter values ​​for completing the query model training.

[0038] As an optional embodiment, after receiving the voice query information and before obtaining the feature information of the image to be identified, the image to be identified corresponding to the voice query information is obtained, wherein the image to be identified is photographed after receiving the voice query information or obtained from a web page after receiving the voice query information.

[0039] The moment of obtaining the image to be recognized corresponding to the voice query information can be after receiving the voice query and before obtaining the feature information of the image to be recognized. For example, after receiving the voice query information, the image to be recognized is obtained, and then the feature information of the image to be recognized is obtained.

[0040] The voice query method of an embodiment of the present invention can be used as an automatic question-and-answer method for the blind based on image analysis. In an optional application scenario, a blind person walking and wanting to understand the surrounding environment can ask the question "Are we in a park now?" After receiving the voice question, a result is obtained based on the feature information in the collected image, and the voice result "Yes" is output. If a user asks, "What's on the ground?" After receiving the user's voice question, a result is obtained based on the feature information in the collected image, and the voice result "Pigeons and bicycles" is output. In an optional application scenario, when a blind user is surfing the Internet, the user's voice query information can be received, for example, "What's in the picture?" Feature information of the image on the web page is obtained, which may be that there are leaves in the picture. After finding image information that matches the user's question, the voice result "Leaves" is played. This voice query method can improve the user's Internet experience and enhance the flexibility of blind people's lives.

[0041] Through the technical solutions of the embodiments of the present invention, more questions and answers related to the visual environment can be provided to the blind group, allowing the blind to better understand the surrounding environment or information on the web page, rather than just staying at the text recognition stage.

[0042] The present invention also provides a preferred embodiment, which includes the following parts:

[0043] The technical solution of the present invention can be implemented through a voice query system. The internal principle of the system mainly adopts the principle of an image-based voice query system, which adopts a method of combining image recognition with natural language processing and is a multi-model learning problem.

[0044] Since the system needs to give relevant answers about the content of the visual scene, object recognition is a must. Secondly, since it needs to process questions raised by users, text analysis is also a must.

[0045] Figure 3 FIG. 1 is a schematic diagram showing the principle of a voice query system according to an embodiment of the present invention. Figure 3 As shown, the voice query process performed by the voice query system can include a training phase and a testing phase. The training phase includes: image acquisition, which involves acquiring images for training; performing feature extraction on the acquired images; and inputting the extracted features into an encoder for encoding. Similarly, a training question set is also input into the encoder, which generates a model based on the training images and the training question set. In the testing phase, a test question is input into the generated model to generate a decoder, which can output the answer to the question, thereby completing model training.

[0046] Specifically, the voice query system of the embodiment of the present invention includes the following modules:

[0047] 1. Object Feature Extraction Module

[0048] The ResNet convolutional neural network architecture is used in the object recognition module, and then fine-tuned on the ImageNet dataset (22,000 categories), which can well cover the common object categories in life. The high-level features of the objects in the picture extracted by Restnet are used (because it is currently the network framework that can achieve the best results in the field of object classification). First, the input size of the image (the resolution of the camera) is obtained as 1280*720, then it is scaled to 448*448 and used as the input of ResNet. Finally, the feature data of the last pooling layer of ResNet is extracted, with a dimension of 1024*14*14, denoted as I. 14*14 corresponds to the number of regions in the input image, and 1024 is the feature dimension corresponding to each region. The size of these regions corresponds to 32*32 in the original image. In addition to the above-mentioned size values, other values ​​can also be used. This is only used as an example and is not used to limit specific values.

[0049] 2. Problem Model Construction Module

[0050] The latest research shows that both long short-term memory neural networks (LSTM) and convolutional neural networks (CNN) have good capabilities for capturing text semantics, but the embodiments of the present invention prefer to use the CNN method.

[0051] First, use the word2vec tool to segment the input question and then vectorize it (convert the language information into the feature information of the text). Then, the vectors of all the words in the question (v i , i=0,..N) into a new vector V I , and finally input the vector into a new CNN network. The CNN used in this module is much smaller than the previous module, generally 3 or so convolution layers are sufficient, and the corresponding convolution kernel sizes are size = 1 (unigram), size = 2 (bigram), and size = 3 (trigram). After the convolution layer processing, a max pooling layer is applied to the output of all the convolution layers to obtain (h1, h2, h3), and then the obtained features are combined to obtain the feature vector H of the problem. c .

[0052] 3. Image and corresponding semantic association module

[0053] The above two modules obtain the features I of each region in the image and the features H of the question c, again using Stacked Attention Networks to predict the answer through multiple inferences.

[0054] Typically, an answer is only relevant to an object in one region of an image, but images typically contain many objects. Therefore, we leverage global image features to predict the answer. Noise from areas unrelated to the answer can lead to a number of possible answers. By applying multiple attention layers, irrelevant noise regions are gradually filtered out, making the answer more relevant to the question.

[0055] First, the features I and H c Perform one-step neural network layer processing, and then use the softmax layer to generate the attention distribution vector P of all regions in the original image, that is, the response probability of each region to the question, and finally calculate the weighted feature weight of each region of the image I P I P With H c Then input it into the stacked coding network to obtain encoder C.

[0056] When a user asks a question, the question is first vectorized, and then the text features are input into the decoder to get the answer to the question.

[0057] The voice query method of the embodiments of the present invention can be an application-based product. This type of product can be implemented as follows: After a blind user installs this application on a terminal (e.g., a mobile phone or computer) while surfing the Internet, they can, in addition to traditional screen readers, ask questions such as "What's in the picture?", "Where is this picture?", "What color is this object?", etc. Upon receiving the user's question, the query results are obtained based on the image features displayed on the interface, and the query results are output as voice output. These are solutions that traditional screen readers cannot provide.

[0058] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0059] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0060] Example 2

[0061] According to an embodiment of the present invention, a voice query device for implementing the above-mentioned voice query method is also provided. Figure 4 is a schematic diagram of an optional voice query device according to an embodiment of the present invention, such as Figure 4 As shown, the device may include:

[0062] The receiving unit 10 is configured to receive voice query information, wherein the voice query information is used to indicate a query keyword.

[0063] The acquisition unit 20 is configured to acquire feature information of the image to be identified, wherein the feature information is used to represent an object in the image to be identified.

[0064] The query unit 30 is configured to query the feature information for target feature information that matches the query keyword.

[0065] The playing unit 40 is configured to play a voice query result when the target feature information is found, wherein the voice query result is used to indicate a target object in the image, and the target object is represented by the target feature information.

[0066] It should be noted that the receiving unit 10 in this embodiment can be used to execute step S202 in embodiment 1 of the present application, the acquisition unit 20 in this embodiment can be used to execute step S204 in embodiment 1 of the present application, the query unit 30 in this embodiment can be used to execute step S206 in embodiment 1 of the present application, and the playback unit 40 in this embodiment can be used to execute step S208 in embodiment 1 of the present application.

[0067] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiment 1. It should be noted that the above modules as part of the device can be run in Figure 1 In the hardware environment shown, it can be implemented by software or by hardware.

[0068] Through the above modules, the technical problem of the voice query system being unintelligent due to the inability to recognize image feature information can be solved, thereby achieving the technical effect of improving the intelligence level of the voice query system.

[0069] Optionally, the query unit 30 includes: an input module for inputting feature information and query keywords into a trained query model, wherein the trained query model is used to query target feature information matching the query keywords from the feature information; and a determination module for determining that the target feature information has been queried when the query model outputs the target feature information.

[0070] Optionally, the device also includes: a training unit, which is used to train a preset query model through a convolutional neural network before querying target feature information matching the query keyword from the feature information to obtain a trained query model, wherein, during the training process, the object features in the pre-obtained multiple images and the pre-obtained information features are used as inputs of the query model, the object features are used to represent the objects in the multiple images, and the information features are used to represent the query questions in a predetermined query question set.

[0071] An embodiment of the present invention further provides a voice query device for implementing the above-mentioned voice query method. Figure 5 is a schematic diagram of an optional voice query device according to an embodiment of the present invention, such as Figure 5 As shown, the device may include:

[0072] The voice interaction module 110 is configured to receive voice query information, wherein the voice query information is used to indicate a query keyword.

[0073] The camera 120 is used to capture the image to be recognized.

[0074] The processor 130 (not shown) is used to obtain feature information of the image to be identified, and search the feature information for target feature information matching the query keyword, wherein the feature information is used to represent the object in the image to be identified.

[0075] The speaker 140 (not shown in the figure) is used to play the voice query result when the target feature information is queried, wherein the voice query result is used to indicate the target object in the image, and the target object is represented by the target feature information.

[0076] Optionally, the device further includes a wireless transceiver 150 for communicating with a preset server. The wireless transceiver 50 can be connected to a backend server. For example, communication with the backend server (preset server) can be achieved via a wireless network. The device can send captured images to be recognized to the backend server, send received voice query information to the backend server, and receive voice query results generated by the server. The preset server can be a cloud server or a more powerful mobile terminal.

[0077] An embodiment of the present invention further provides a head-mounted device for implementing the above-mentioned voice query method. Figure 7 is a schematic diagram of another optional head-mounted device according to an embodiment of the present invention, such as Figure 7 As shown, the device may include:

[0078] The voice interaction module 702 is configured to receive voice query information of a currently identified object, wherein the currently identified object is wearing a head-mounted device, and the voice query information carries query keywords for searching for a target environment where the currently identified object is located;

[0079] Camera 704, used to capture an image of the environment observed by the current identification object at the current observation angle;

[0080] Processor 706, configured to obtain feature information of the environment image and search the feature information for target feature information matching the query keyword, wherein the feature information is used to represent an object in the environment image;

[0081] Speaker 708 is used to play the voice query result when the target feature information is found, wherein the voice query result is used to indicate the target object in the environment image, and the target object is represented by the target feature information.

[0082] The voice query device of the embodiment of the present invention can be a mobile device type product, for example, it can be a pair of glasses, a hat, or other devices worn on the user. The type of product form can be: a simple device that can be worn on the head, which is equipped with three miniature high-definition cameras and can obtain 360-degree images of the surrounding scene. A network signal transceiver (wireless transceiver) is configured on the top to transmit data with the background. The hardware configuration here needs to be lightweight and safe. The questioning method uses voice interaction, and the answer response given by the system also uses voice.

[0083] As an optional method, the processor is further configured to:

[0084] S1, inputting feature information and query keywords into a trained query model, wherein the trained query model is used to search for target feature information matching the query keywords from the feature information;

[0085] S2: When the query model outputs the target feature information, it is determined that the target feature information is found.

[0086] As an optional manner, the above-mentioned head-mounted device also includes: a wireless transceiver for communicating with a preset server.

[0087] As an optional embodiment, the wireless transceiver includes:

[0088] A networking module, configured to establish a network connection with a preset server in response to a network connection request;

[0089] The transceiver module is used to send the environment image and voice query information to the preset server when a network connection is established with the preset server; receive the background recognition results returned by the preset server; and send the background recognition results to the speaker to play the background recognition results.

[0090] In an optional application scenario, when a blind user is surfing the Internet, a network connection request from the user can be received to establish a network connection with a preset server; further, when the user's voice query information is received, the voice query information and the acquired environmental image are sent to the preset server through the network for background recognition in the preset server. It can be understood that the computing power of the query model running in the preset server is higher than that of the query model preset in the above-mentioned head-mounted device, and the recognition result is more accurate. For example, the user's question is "What's in the picture?" and the voice query information and the picture are sent to the preset server through the transceiver module. The feature information of the image is obtained in the preset server. The feature information may be that there are leaves in the picture. After the picture information matching the user's question is queried in the preset server, the background recognition result returned by the preset server is received, and the voice result "leaves" is played. Through such a voice query method, the user's Internet experience can be improved and the flexibility of the blind person's life can be improved.

[0091] As an optional method, after receiving the voice query information and before obtaining the feature information of the image to be identified, the device is also used to: obtain the image to be identified corresponding to the voice query information, wherein the image to be identified is photographed after receiving the voice query information, or obtained from a web page after receiving the voice query information.

[0092] As an optional manner, after receiving the voice query information, the apparatus is further configured to: when the voice query information is query information of an environmental awareness type, determine that the current identification object is an identification object of a visual impairment type.

[0093] It is understood that when the voice query information received by the head-mounted device is of an environmental awareness type, the head-mounted device can determine that the current user is a blind user based on the type of the received query information. For example, if the received query information is a question used to understand objective things in the environment, such as "What's on the ground?" or "Is this a park?", the head-mounted device can determine that the current user is a blind user.

[0094] As an optional method, the trained query model running in the above-mentioned processor is trained in the following manner: a pre-set query model is trained through a convolutional neural network to obtain a trained query model, wherein, during the training process, object features in multiple images obtained in advance and information features obtained in advance are used as inputs of the query model, the object features are used to represent objects in multiple images, and the information features are used to represent query questions in a predetermined query question set.

[0095] As an optional method, the pre-set query model is trained by the convolutional neural network to obtain a trained query model, including:

[0096] S1, obtain the correlation between object features and information features;

[0097] S2, continuously adjusting the values ​​of the parameters in the query model until the highest correlation is obtained, wherein the values ​​of the parameters in the trained query model are the values ​​of the parameters when the correlation is the highest.

[0098] It is understood that, when the trained query model is obtained through the training method, the query model can be pre-installed in the processor. In another optional embodiment, when the trained query model is obtained through the training method, the trained query model is received through the wireless transceiver to run the query model in the processor.

[0099] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiment 1. It should be noted that the above modules as part of the device can be run in Figure 1 The hardware environment shown can be implemented through software or hardware, wherein the hardware environment includes a network environment.

[0100] Example 3

[0101] According to an embodiment of the present invention, a server or terminal for implementing the above-mentioned voice query method is also provided.

[0102] Figure 6is a structural block diagram of a terminal according to an embodiment of the present invention, such as Figure 6 As shown, the terminal may include: one or more (only one is shown in the figure) processors 201, a memory 203, and a transmission device 205 (such as the sending device in the above embodiment), as shown in FIG. Figure 6 As shown, the terminal may further include an input and output device 207 .

[0103] Among them, the memory 203 can be used to store software programs and modules, such as the program instructions / modules corresponding to the voice query method and device in the embodiments of the present invention. The processor 201 executes various functional applications and data processing by running the software programs and modules stored in the memory 203, that is, implementing the above-mentioned voice query method. The memory 203 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 203 may further include a memory remotely located relative to the processor 201, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0104] The transmission device 205 is used to receive or send data via a network, and can also be used for data transmission between a processor and a memory. Specific examples of the network may include wired networks and wireless networks. In one embodiment, the transmission device 205 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 205 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0105] Specifically, the memory 203 is used to store application programs.

[0106] The processor 201 can call the application stored in the memory 203 through the transmission device 205 to perform the following steps: receiving voice query information, wherein the voice query information is used to indicate the query keyword; obtaining feature information of the image to be identified, wherein the feature information is used to represent the object in the image to be identified; searching for target feature information matching the query keyword from the feature information; and when the target feature information is found, playing the voice query result, wherein the voice query result is used to indicate the target object in the image, and the target object is represented by the target feature information.

[0107] The processor 201 is also used to perform the following steps: input the feature information and query keywords into the trained query model, wherein the trained query model is used to query the target feature information that matches the query keyword from the feature information; when the query model outputs the target feature information, it is determined that the target feature information has been queried.

[0108] The processor 201 is also used to perform the following steps: training a preset query model through a convolutional neural network to obtain a trained query model, wherein, during the training process, the object features in the pre-obtained multiple images and the pre-obtained information features are used as inputs of the query model, the object features are used to represent the objects in the multiple images, and the information features are used to represent the query questions in the predetermined query question set.

[0109] Processor 201 is also used to perform the following steps: obtaining the correlation between object features and information features; continuously adjusting the values ​​of parameters in the query model until the highest correlation is obtained, wherein the values ​​of parameters in the trained query model are the values ​​of parameters when the correlation is the highest.

[0110] The processor 201 is further configured to execute the following steps: obtaining an image to be identified corresponding to the voice query information, wherein the image to be identified is photographed after receiving the voice query information, or is obtained from a web page after receiving the voice query information.

[0111] By adopting the embodiment of the present invention, by receiving voice query information and obtaining feature information of the image to be recognized, and then searching for target feature information matching the query keyword from the feature information, and playing the voice query result when the target feature information is found, the purpose of outputting the query result by voice after recognizing the feature information in the image is achieved, thereby achieving the technical effect of improving the intelligence of the voice query system, and further solving the technical problem of the voice query system being unintelligent due to the inability to recognize image feature information.

[0112] Optionally, the specific examples in this embodiment may refer to the examples described in the above-mentioned embodiment 1 and embodiment 2, and this embodiment will not be described in detail here.

[0113] It can be understood by those skilled in the art that Figure 6 The structure shown is for illustration only, and the terminal may be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Figure 6 It does not limit the structure of the above electronic device. For example, the terminal may also include Figure 6More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 6 Different configurations shown.

[0114] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0115] Example 4

[0116] The embodiment of the present invention further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to execute the program code of the voice query method.

[0117] Optionally, in this embodiment, the above-mentioned storage medium may be located on at least one network device among the multiple network devices in the network shown in the above-mentioned embodiment.

[0118] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps:

[0119] S1, receiving voice query information, wherein the voice query information is used to indicate a query keyword;

[0120] S2, obtaining feature information of the image to be identified, wherein the feature information is used to represent the object in the image to be identified;

[0121] S3, searching the feature information for target feature information that matches the query keyword;

[0122] S4: When the target feature information is found, a voice query result is played, wherein the voice query result is used to indicate the target object in the image, and the target object is represented by the target feature information.

[0123] Optionally, the storage medium is also configured to store program code for executing the following steps: inputting feature information and query keywords into a trained query model, wherein the trained query model is used to query target feature information that matches the query keywords from the feature information; and when the query model outputs the target feature information, determining that the target feature information has been queried.

[0124] Optionally, the storage medium is also configured to store program code for executing the following steps: training a preset query model through a convolutional neural network to obtain a trained query model, wherein, during the training process, object features in multiple images obtained in advance and information features obtained in advance are used as inputs to the query model, the object features are used to represent objects in multiple images, and the information features are used to represent query questions in a predetermined query question set.

[0125] Optionally, the storage medium is also configured to store program code for performing the following steps: obtaining the correlation between object features and information features; continuously adjusting the values ​​of parameters in the query model until the highest correlation is obtained, wherein the values ​​of the parameters in the trained query model are the values ​​of the parameters when the correlation is the highest.

[0126] Optionally, the storage medium is further configured to store program code for executing the following steps: obtaining an image to be identified corresponding to the voice query information, wherein the image to be identified is photographed after receiving the voice query information, or is obtained from a web page after receiving the voice query information.

[0127] Optionally, the specific examples in this embodiment may refer to the examples described in the above-mentioned embodiment 1 and embodiment 2, and this embodiment will not be described in detail here.

[0128] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.

[0129] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0130] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0131] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0132] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.

[0133] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0134] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0135] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A head-mounted device, characterized in that: include: a voice interaction module, configured to receive voice query information of a currently identified object, wherein the currently identified object is wearing the head-mounted device, and the voice query information carries query keywords for querying a target environment where the currently identified object is located; A camera, configured to capture an image of the environment under a current viewing angle of the currently identified object after receiving the voice query information; a processor, configured to obtain feature information of the environment image, and search the feature information for target feature information matching the query keyword, wherein the feature information is used to represent an object in the environment image; A loudspeaker is used to play a voice query result when the target feature information is found, wherein the voice query result is used to indicate a target object in the environment image, and the target object is represented by the target feature information.

2. The device according to claim 1, characterized in that The processor is further configured to: Inputting the feature information and the query keyword into a trained query model, wherein the trained query model is used to query the target feature information matching the query keyword from the feature information; In a case where the query model outputs the target feature information, it is determined that the target feature information is found in the query.

3. The device according to claim 1, characterized in that The head-mounted device further comprises: A wireless transceiver is used to communicate with a preset server.

4. The device according to claim 3, characterized in that The wireless transceiver comprises: A networking module, configured to establish a network connection with a preset server in response to a network connection request; The transceiver module is used to send the environmental image and the voice query information to the preset server when a network connection is established with the preset server; receive the background recognition result returned by the preset server; and send the background recognition result to the speaker to play the background recognition result.

5. The device according to claim 1, characterized in that After receiving the voice query information and before acquiring the feature information of the image to be recognized, the device is further configured to: Acquire the image to be recognized corresponding to the voice query information, wherein the image to be recognized is acquired from a web page after receiving the voice query information; The processor is further configured to obtain feature information of the image to be identified, and to search the feature information for target feature information that matches the query keyword, wherein the feature information is used to represent an object in the image to be identified.

6. The device according to claim 1, characterized in that After receiving the voice query information, the device is further configured to: In a case where the voice query information is query information of an environmental awareness type, it is determined that the current identified object is an identified object of a visual impairment type.

7. The device according to claim 2, characterized in that The trained query model running in the processor is trained in the following manner: A pre-set query model is trained through a convolutional neural network to obtain a trained query model, wherein, during the training process, object features in a plurality of pre-obtained images and pre-obtained information features are used as inputs to the query model, the object features are used to represent objects in the plurality of images, and the information features are used to represent query questions in a predetermined query question set.

8. The device according to claim 7, characterized in that The method of training a preset query model by a convolutional neural network to obtain a trained query model includes: Obtaining the correlation between the object feature and the information feature; The values ​​of the parameters in the query model are continuously adjusted until the highest relevance is obtained, wherein the values ​​of the parameters in the trained query model are the values ​​of the parameters when the relevance is the highest.

Citation Information

Patent Citations

  • Wearable guide apparatus for totally blind people

    CN104473717A

  • Keyword notification method, equipment and computer program product based on character recognition

    CN105518712A