Intelligent object searching method, device and equipment and intelligent sound box
By obtaining user voice information, identifying the wake-up word, and calling a multi-dimensional object feature library, the problems of low accuracy and efficiency in object search in existing technologies are solved, and efficient object positioning and guidance of smart speakers are achieved.
Patent Information
- Application Number
- CN202510602828.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The accuracy and efficiency of finding objects in existing technologies are low, and manually recording the location of objects is time-consuming and prone to errors.
By obtaining user voice information, recognizing voice wake-up words and generating deep object-finding language content, calling the multi-dimensional item feature library to determine the location of the item, and using smart speakers to guide users to find the item.
It improves the accuracy and efficiency of finding objects and enhances the interaction between users and smart speakers.
Smart Images

Figure CN120612933A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to intelligent object-finding methods, devices, equipment, and smart speakers. Background Art
[0002] As people's lives become more and more affluent, they buy a lot of items. In order to keep the indoor space tidy, they usually store items in different spaces, such as drawers, storage cabinets, etc. However, in the long-term production and life, most items will be forgotten, resulting in users being unable to find them when they need them. At present, the common way to find items is to manually record the specific location information when storing items. However, the number of objects is huge, and a single manual record will waste a lot of time. In addition, manual recording can only record the approximate location, which is prone to deviations when searching. Therefore, the accuracy and efficiency of the above method of finding items are low.
[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide an intelligent object-finding method, device, equipment and smart speaker, aiming to solve the technical problems of low accuracy and efficiency of object-finding in the existing technology.
[0005] To achieve the above objectives, this application proposes an intelligent object-finding method, which includes:
[0006] Obtaining the user's current voice information and determining a voice wake-up word based on the current voice information;
[0007] When the voice wake-up word is a find-object wake-up word, generating deep find-object language content according to the current voice information;
[0008] Acquire a multi-dimensional item feature library of the target area by calling a resource interface, and determine the location information of the item corresponding to the deep object-finding language content based on the multi-dimensional item feature library;
[0009] Targeted object-finding guidance information is generated according to the location information of the object, and the smart speaker guides the user to find the object according to the targeted object-finding guidance information.
[0010] In one embodiment, the step of obtaining the user's current voice information and determining the voice wake-up word based on the current voice information includes:
[0011] Acquire the user's current voice information based on the multi-channel voice pickup strategy;
[0012] Performing sound wave recognition on the current voice information to obtain current sound wave information of the user;
[0013] Denoising the current sound wave information according to a target noise reduction algorithm, and extracting sound wave feature information of the denoised sound wave information;
[0014] The acoustic wave feature information is input into a target wake-up word recognition model, and the voice wake-up word recognized and output by the target wake-up word recognition model is obtained.
[0015] In one embodiment, when the voice wake-up word is a find-object wake-up word, the step of generating deep find-object language content according to the current voice information includes:
[0016] When the voice wake-up word is a find-and-seek wake-up word, performing data cleaning on the current voice information;
[0017] Filtering the cleaned voice information and performing information expansion on the filtered voice information;
[0018] generating target voice information according to the filtered voice information and the expanded voice information;
[0019] Converting the target voice information to obtain target text information;
[0020] The target text information is subjected to deep reasoning according to a natural language processing strategy to obtain deep object-seeking language content.
[0021] In one embodiment, when the voice wake-up word is a find-object wake-up word, the step of generating deep find-object language content according to the current voice information includes:
[0022] When the voice wake-up word is a find-and-seek wake-up word, performing data cleaning on the current voice information;
[0023] Filtering the cleaned voice information and performing information expansion on the filtered voice information;
[0024] generating target voice information according to the filtered voice information and the expanded voice information;
[0025] Converting the target voice information to obtain target text information;
[0026] The target text information is subjected to deep reasoning according to a natural language processing strategy to obtain deep object-seeking language content.
[0027] In one embodiment, the step of determining the location information of the object corresponding to the deep object search language content based on the predicted probability, the attribute information, and the multi-dimensional object feature library includes:
[0028] Traversing the same desired items between each of the desired item sets, and counting the number of the same desired items;
[0029] When the number is greater than a preset value, selecting items corresponding to the deep object-seeking language content from the same intended items whose number is greater than the preset value according to the predicted probability and the attribute information;
[0030] The location information of the object is determined according to the multi-dimensional object feature library.
[0031] In one embodiment, the multi-dimensional item feature library includes a spatial feature database, an item video image feature library, and an item radar image feature library;
[0032] The step of determining the location information of the item based on the multi-dimensional item feature library includes:
[0033] Obtaining a spatial item feature information set from the spatial feature database, and matching the feature information of the item with each feature information in the spatial item feature information set;
[0034] Determine the spatial information where the item is located based on the spatial item feature information matching result;
[0035] filtering the object video image feature library and the object radar image feature library respectively according to the spatial information;
[0036] Acquire a set of item video feature information from the filtered item video image feature library, and match the feature information of the item with each feature information in the item video feature information;
[0037] Determining first location information of the object based on a matching result of the object video feature information;
[0038] Obtaining a set of item radar feature information from the filtered item radar image feature library, and matching the feature information of the item with each feature information in the set of item radar feature information;
[0039] The second location information of the object is determined according to the matching result of the radar feature information of the object, and the location information of the object is determined according to the first location information and the second location information.
[0040] In one embodiment, before the step of acquiring the multi-dimensional item feature library of the target area by calling the resource interface, the method further includes:
[0041] If it is detected that the user has not imported the floor plan within the preset time period, a target floor plan is generated based on the user's location information;
[0042] Performing two-dimensional recognition on the target floor plan, assigning three-dimensional parameters to the recognized plane space data, and obtaining the current three-dimensional space of the target area;
[0043] Building a spatial feature database based on feature information of each object in the current three-dimensional space and the current three-dimensional space;
[0044] Controlling the camera device to collect video image data of each object in the current three-dimensional space, and determining first position information of each object based on video feature information of the object in the video image data;
[0045] Building an object video image feature library based on the object video feature information and each of the first position information;
[0046] Controlling the radar to collect radar image data of each object in the current three-dimensional space, and determining second position information of each object based on radar feature information of the object in the radar image data;
[0047] Building an object radar image feature library based on the object radar feature information and each of the second position information;
[0048] A multi-dimensional item feature library of the target area is generated according to the spatial feature database, the item video image feature library and the item radar image feature library, and the multi-dimensional item features of the target area are uploaded.
[0049] In addition, to achieve the above-mentioned purpose, the present application also proposes an intelligent object-finding device, which includes:
[0050] A determination module, configured to obtain the user's current voice information and determine a voice wake-up word based on the current voice information;
[0051] A generation module, configured to generate deep object-finding language content based on the current voice information when the voice wake-up word is a find-object wake-up word;
[0052] The determination module is further configured to obtain a multi-dimensional item feature library of the target area by calling a resource interface, and determine the location information of the item corresponding to the deep object search language content based on the multi-dimensional item feature library;
[0053] The search module is used to generate target object-finding guidance information according to the location information of the object, and guide the user to find the object according to the target object-finding guidance information based on the smart speaker.
[0054] In addition, to achieve the above-mentioned purpose, the present application also proposes an intelligent object-finding device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the intelligent object-finding method as described above.
[0055] In addition, to achieve the above-mentioned purpose, the present application also proposes a smart speaker, which includes: a speaker, an indicator, a display, a memory, a processor, and a computer program stored on the memory and runnable on the processor. When the computer program is executed by the processor, the steps of the smart object-finding method described above are implemented.
[0056] One or more technical solutions proposed in this application have at least the following technical effects: by obtaining the user's current voice information and determining a voice wake-up word based on the current voice information; when the voice wake-up word is a find-and-seek wake-up word, generating deep find-and-seek language content based on the current voice information; by calling a resource interface to obtain a multi-dimensional item feature library of the target area, and determining the location information of the item corresponding to the deep find-and-seek language content based on the multi-dimensional item feature library; generating target find-and-seek guidance information based on the location information of the item, and guiding the user to find the item based on the target find-and-seek guidance information based on the smart speaker. Through the above method, after determining the voice wake-up word based on the user's current voice information, it is determined whether the voice wake-up word is a find-and-seek wake-up word. If so, it indicates that the user needs to find the item. At this time, the complex current voice information is deeply analyzed in combination with the context extension information, and then the location information of the item is determined by calling the multi-dimensional item feature library in the spatial dimension, video image dimension, and radar image dimension. At this time, the smart speaker is used to instruct the user to find the item, thereby effectively improving the accuracy and efficiency of find-and-seek and enhancing the interaction between the user and the smart speaker. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0058] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0059] Figure 1 A flowchart of the first embodiment of the intelligent object-finding method of this application is provided;
[0060] Figure 2A schematic diagram of the process of constructing a spatial feature database provided in Example 1 of the intelligent object-finding method of this application;
[0061] Figure 3 A flowchart of the second embodiment of the intelligent object-finding method of this application is provided;
[0062] Figure 4 This is a schematic diagram of the module structure of the intelligent object-finding device according to an embodiment of the present application;
[0063] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the intelligent object-finding method in the embodiment of the present application.
[0064] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0065] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, such as an intelligent object-finding device. The following uses an intelligent object-finding device as an example to illustrate this embodiment and the following embodiments.
[0066] Based on this, the embodiment of the present application provides an intelligent object-finding method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the intelligent object-finding method of this application.
[0067] In this embodiment, the intelligent object-finding method includes steps S10 to S40:
[0068] Step S10: Acquire the user's current voice information and determine a voice wake-up word based on the current voice information.
[0069] It should be noted that the current voice information refers to the voice information collected when a sound is emitted in the user's target area, and the voice wake-up word refers to the keyword used to wake up the smart device using voice information. The voice wake-up word can be a search wake-up word, for example, looking for tools for rust removal. The voice wake-up word can also be a start wake-up word, for example, starting a sweeping robot.
[0070] Furthermore, step S10 includes: obtaining the user's current voice information according to the multi-channel sound pickup strategy; performing sound wave recognition on the current voice information to obtain the user's current sound wave information; performing noise reduction on the current sound wave information according to the target noise reduction algorithm, and extracting the sound wave feature information of the noise-reduced sound wave information; inputting the sound wave feature information into the target wake-up word recognition model, and obtaining the voice wake-up word recognized and output by the target wake-up word recognition model.
[0071] It should be understood that in order to effectively improve the accuracy of picking up the user's current voice information, a multi-channel pickup strategy is adopted to obtain the user's current voice information, wherein the multi-channel pickup strategy refers to the simultaneous collection of voice information in different positions or directions through multiple independent audio collection channels. In addition, in order to effectively improve the efficiency of recognizing voice wake-up words, before recognition, it is necessary to perform voiceprint recognition on the current voice information, perform noise reduction on the current sound wave information, and extract sound wave feature information. wherein, performing noise reduction on the current sound wave information according to the target noise reduction algorithm can effectively filter out the environmental noise in the current sound wave information. The target wake-up word recognition model refers to a model for identifying the voice wake-up words corresponding to the voiceprint information. After obtaining the sound wave feature information, the sound wave feature information is input into the target wake-up word recognition model. At this time, the target wake-up word recognition model will recognize and output the voice wake-up words corresponding to the sound wave feature information.
[0072] Step S20: When the voice wake-up word is a find-object wake-up word, deep find-object language content is generated according to the current voice information.
[0073] It can be understood that deep object search language content refers to the voice content obtained by deeply analyzing the complex current voice information in combination with contextual extension information. After obtaining the voice wake-up word corresponding to the sound wave feature information, it is necessary to determine whether the voice wake-up word is a search wake-up word. If so, it indicates that the user needs to find an item. At this time, deep object search language content is generated based on the current voice information.
[0074] Step S30 , obtaining a multi-dimensional object feature library of the target area by calling a resource interface, and determining the location information of the object corresponding to the deep object search language content according to the multi-dimensional object feature library.
[0075] It should be understood that the resource interface refers to the interface for obtaining the multi-dimensional item feature library of the target area. Calling the resource interface can obtain the multi-dimensional item feature library of the target area stored in the local unit, or the multi-dimensional item feature library of the target area stored in the cloud. The specific method can be determined according to the network quality. For example, when the network quality is better, the multi-dimensional item feature library of the target area stored in the cloud is obtained first. Conversely, when the network quality is poor, the multi-dimensional item feature library of the target area stored in the local unit is obtained first. The target area can be the indoor area where the user is located. The multi-dimensional item feature library includes a spatial feature database, an item video image feature library, and an item radar image feature library, that is, the location information of the item corresponding to the deep object search language content is comprehensively determined from the spatial dimension, video image dimension, and radar image dimension, thereby effectively improving the accuracy of determining the location information of the item.
[0076] Furthermore, before the step of obtaining the multi-dimensional item feature library of the target area by calling the resource interface, it also includes: if it is detected that the user has not imported the floor plan within a preset time period, generating a target floor plan based on the user's location information; performing two-dimensional recognition on the target floor plan, assigning three-dimensional parameters to the recognized plane space data, and obtaining the current three-dimensional space of the target area; constructing a spatial feature database based on the feature information of each item in the current three-dimensional space and the current three-dimensional space; controlling the camera device to collect video image data of each item in the current three-dimensional space, and determining the first position information of each item based on the item video feature information of the video image data; constructing an item video image feature library based on the item video feature information and each first position information; controlling the radar to collect radar image data of each item in the current three-dimensional space, and determining the second position information of each item based on the item radar feature information of the radar image data; constructing an item radar image feature library based on the item radar feature information and each second position information; generating a multi-dimensional item feature library of the target area based on the spatial feature database, the item video image feature library, and the item radar image feature library, and uploading the multi-dimensional item features of the target area.
[0077] It is understandable that the multi-dimensional item feature library of the target area can be pre-built according to the floor plan before the intelligent object search is carried out. In this embodiment, an interface can be provided for the user to import the floor plan. The floor plan can be drawn by the user or directly generated by a third-party application. This embodiment does not limit this. If it is detected that the user has not imported the floor plan within the preset time period, it indicates that the floor plan needs to be generated. At this time, the target floor plan of the target area is generated according to the user's location information. The target floor plan includes spaces such as the master bedroom, second bedroom, living room, and bathroom, and each space can be numbered. Click on the specific number to enter the three-dimensional interface of the corresponding space. Reference Figure 2 , Figure 2 A flowchart for constructing a spatial feature database is provided, specifically: obtaining a target floor plan of a target area imported by a user or generating a target floor plan based on the user's location information; performing two-dimensional recognition on the target floor plan of the target area to obtain plane space data; assigning three-dimensional parameters to the plane space data to obtain the current three-dimensional space of the target area; and constructing a spatial feature database based on the characteristic information of each object in the current three-dimensional space and the current three-dimensional space.
[0078] It should be noted that the current 3D space can be a relatively independent single indoor space. For example, the target area is a residential area, which can be a study, bedroom, living room, and kitchen. The indoor space can be further divided into one or more spatial modes, based on a classification logic that divides the indoor space into open space and closed space. This classification logic can be open space = directly exposed and freely accessible; closed space = requires opening or unlocking, hidden storage. For example, open space includes but is not limited to walls, ceilings, floors, desktops, and countertops, while closed space includes but is not limited to drawers, wardrobes, cabinets, TV cabinets, under-bed storage boxes, and covered storage boxes. After determining the current 3D space of the target area, a spatial feature database is constructed by combining the feature information of each item within the current 3D space.
[0079] It should be understood that the object video image feature library can be constructed by video image data of each object in the current three-dimensional space collected by a camera device. The camera device can be at least one of a color camera, a depth camera, and a dynamic vision (DVS) camera. The specific construction process is: after collecting the video image data of each object in the open space and the closed space, the position of each object will be identified, which is the first position information. The first position information represents the position information determined from the video image dimension. In addition, the video image data of each object will be pre-processed by data cleaning, feature extraction and other pre-processing operations to obtain the object video feature information of the video image data. At this time, the object video image feature library can be constructed in combination with the first position information of each object.
[0080] It should be noted that, similarly, the object radar image feature library can be constructed by collecting radar image data of each object in the current three-dimensional space. The radar image data includes but is not limited to millimeter wave radar image data, ultra-wideband radar (UWB) image data, etc. The specific construction process is: after collecting the radar image data of each object in the current three-dimensional space, the position of each object will be identified, which is the second position information, wherein the second position information represents the position information determined from the radar image dimension. In addition, the radar image data of each object will be pre-processed by data cleaning, feature extraction and other pre-processing operations to obtain the object radar feature information of the radar image data. At this time, the object radar image feature library can be constructed in combination with the second position information of each object.
[0081] Furthermore, the step of determining the location information of the items corresponding to the deep object search language content based on the multi-dimensional item feature library includes: performing keyword recognition on the deep object search language content to obtain each deep object search keyword; determining the intended item set and predicted probability corresponding to each deep object search keyword based on the item prediction model; obtaining the attribute information of each intended item in the intended item set; and determining the location information of the items corresponding to the deep object search language content based on the predicted probability, the attribute information and the multi-dimensional item feature library.
[0082] It can be understood that the deep object search keywords represent the keywords contained in the deep object search language content. There are many types of deep object search keywords, for example, verb keywords and noun keywords. In order to effectively improve the determination of the intended items corresponding to each deep object search keyword, the item prediction model is connected, and each deep object search keyword is input into the item prediction model. The item prediction model outputs the intended item set and prediction probability corresponding to each deep object search keyword. For example, the intended items corresponding to the deep object search keyword a are A1, A2 and A3. At this time, the intended item set A is composed of the intended items A1, A2 and A3. The attribute information includes but is not limited to the material, shape, size and user of the intended item. After obtaining the predicted probability and attribute information, the location information of the item corresponding to the deep object search language content is determined in combination with the multi-dimensional item feature library.
[0083] Furthermore, the step of determining the location information of the item corresponding to the deep object search language content based on the predicted probability, the attribute information and the multi-dimensional item feature library includes: traversing the same intention items between each of the intended item sets, and counting the number of the same intention items; when the number is greater than a preset value, screening the items corresponding to the deep object search language content from the same intention items whose number is greater than the preset value based on the predicted probability and the attribute information; and determining the location information of the item based on the multi-dimensional item feature library.
[0084] It should be understood that the term "identical intended item" refers to an intended item that exists simultaneously in multiple sets of intended items. Since different deep object search keywords can correspond to the same intended item, after obtaining the identical intended items, a determination is made as to whether the number of identical intended items exceeds a preset value, which may be 1. If so, this indicates the presence of multiple identical intended items. Based on the predicted probability and attribute information, the item that best matches the deep object search language content is selected from the identical intended items whose number exceeds the preset value. For example, the item with the highest predicted probability and the most matching attribute information is selected. The item's location information is then determined from the spatial, video, and radar image dimensions using a multi-dimensional object feature library.
[0085] Furthermore, the multidimensional item feature library includes a spatial feature database, an item video image feature library, and an item radar image feature library; the step of determining the location information of the item based on the multidimensional item feature library includes: obtaining a spatial item feature information set in the spatial feature database, matching the feature information of the item with each feature information in the spatial item feature information set; determining the spatial information where the item is located based on the spatial item feature information matching result; filtering the item video image feature library and the item radar image feature library respectively according to the spatial information; obtaining the item video feature information set in the filtered item video image feature library, matching the feature information of the item with each feature information in the item video feature information; determining the first location information of the item based on the item video feature information matching result; obtaining the item radar feature information set in the filtered item radar image feature library, matching the feature information of the item with each feature information in the item radar feature information set; determining the second location information of the item based on the item radar feature information matching result, and determining the location information of the item based on the first location information and the second location information.
[0086] It is understandable that after obtaining the multi-dimensional item feature library of the target area by calling the resource interface, a convolutional neural network will also be introduced, and the introduction method of the convolutional neural network can be called from the cloud. At this time, feature matching, matching result analysis and feature screening are performed through the convolutional neural network, specifically: matching the feature information of the item with each feature information in the spatial item feature information set, matching the feature information of the item with each feature information in the item video feature information, matching the feature information of the item with each feature information in the item radar feature information set, that is, first determining the spatial information where the item is located, then determining the first position information of the item in the video image dimension, and finally determining the second position information of the item in the radar image dimension. In order to effectively improve the accuracy of determining the location information of the item, this embodiment comprehensively determines the location information of the item from the first position information in the video image dimension and the second position information of the item in the radar image dimension, thereby effectively improving the accuracy of determining the location information.
[0087] Step S40: generating target object-finding guidance information according to the location information of the object, and guiding the user to find the object based on the target object-finding guidance information based on the smart speaker.
[0088] It is understood that the target object guidance information refers to guidance information that guides the user to find the object. The dimensions of the target object guidance information include but are not limited to voice, lighting, and video images. In addition, while the user is searching for the object, the user can also detect the reflected light after the infrared light emitted by the light-emitting diode in the light detector to determine whether the enclosed space is open or closed. For example, the closing and opening of the doors of drawers, wardrobes, and cabinets can be realized by turning off the lights when the door is closed and turning on the lights when the door is opened.
[0089] It should be noted that in order to enhance the interactivity between users and smart homes, this embodiment displays the target object search guidance information based on a smart speaker. The smart speaker includes but is not limited to a speaker, an indicator, and a display. The speaker can play voice-based target object search guidance information, the indicator can display light-based target object search guidance information, and the display can display video-based target object search guidance information. Compared with other mobile terminals, the smart speaker has a single function, and the buttons only include a power button, a volume button, a selection button, a confirmation button, and a cancel button. The above buttons can be mechanical or touch-sensitive, which is more convenient to operate, can effectively improve the accuracy and efficiency of finding objects, and is suitable for a variety of people without becoming addicted. For this reason, this embodiment can also be applied to other smart home devices.
[0090] This embodiment obtains the user's current voice information and determines a voice wake-up word based on the current voice information; when the voice wake-up word is a find-thing wake-up word, generates deep find-thing language content based on the current voice information; obtains a multi-dimensional item feature library of the target area by calling a resource interface, and determines the location information of the item corresponding to the deep find-thing language content based on the multi-dimensional item feature library; generates target find-thing guidance information based on the location information of the item, and guides the user to find the item based on the target find-thing guidance information based on the smart speaker. Through the above method, after determining the voice wake-up word based on the user's current voice information, it is determined whether the voice wake-up word is a find-thing wake-up word. If so, it indicates that the user needs to find something. At this time, the complex current voice information is deeply analyzed in combination with the context extension information, and then the location information of the item is determined by calling the multi-dimensional item feature library of spatial dimension, video image dimension and radar image dimension. At this time, the smart speaker is used to instruct the user to find the item, thereby effectively improving the accuracy and efficiency of find-things and enhancing the interaction between the user and the smart speaker.
[0091] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above embodiment 1 can be referred to the above introduction and will not be described in detail later. Figure 3 Step S20 includes steps S201 to S205:
[0092] Step S201: When the voice wake-up word is a find-thing wake-up word, data cleaning is performed on the current voice information.
[0093] It should be noted that in order to effectively improve the accuracy of the deep object-seeking language content, when determining that the voice wake-up word is the object-seeking wake-up word, it is necessary to perform data cleaning on the current voice information to eliminate invalid information in the current voice information.
[0094] Step S202: filtering the cleaned voice information and performing information expansion on the filtered voice information.
[0095] It is understandable that in order to effectively improve the efficiency of information expansion, after the current voice information is cleaned, it is also filtered to retain the main frequency band of the human voice and filter out ultra-low and ultra-high frequency interference. Because natural language processing strategies require contextual voice information to understand complex filtered voice information during deep reasoning, it is necessary to expand the filtered voice information, and this expanded voice information is strongly correlated with the filtered voice information.
[0096] Step S203: generating target voice information according to the filtered voice information and the expanded voice information.
[0097] It should be understood that after obtaining the expanded voice information, the filtered voice information and the expanded voice information are fused according to the expansion direction of the expanded voice information, and the fused voice information is tested for consistency. When the test passes, the target voice information is obtained.
[0098] Step S204: convert the target voice information to obtain target text information.
[0099] It is understandable that in order to facilitate deep reasoning of the natural language processing strategy, the target voice information needs to be converted into corresponding target text information. At this time, a voice-to-text conversion model trained by a deep learning algorithm can be used, and other text conversion strategies can also be used. This embodiment does not limit this.
[0100] Step S205 , performing deep reasoning on the target text information according to a natural language processing strategy to obtain deep object-finding language content.
[0101] It should be understood that the natural language processing strategy can be a language processing strategy derived from the integration of Transformer's context modeling, data evolution and enhancement, and multi-dimensional text reasoning capabilities. After obtaining the target text information, the natural language processing strategy deeply infers the deep object-seeking language content corresponding to the complex target text information.
[0102] This embodiment performs data cleaning on the current voice information when the voice wake-up word is a find-an-object wake-up word; filters the cleaned voice information, and expands the filtered voice information; generates target voice information based on the filtered voice information and the expanded voice information; converts the target voice information to obtain target text information; and performs deep reasoning on the target text information according to a natural language processing strategy to obtain deep find-an-object language content. Through the above method, when it is determined that the voice wake-up word is a find-an-object wake-up word, data cleaning, filtering and other operations are performed on the current voice information to improve the accuracy of the expanded voice information, and then the filtered voice information and the expanded voice information are fused into the target voice information. After converting the target voice information into the corresponding target text information, the deep find-an-object language content corresponding to the target text information is inferred according to the natural language processing strategy, thereby effectively improving the efficiency of obtaining the deep find-an-object language content.
[0103] This application also provides an intelligent object-finding device, please refer to Figure 4 , the intelligent object-finding device comprises:
[0104] The determination module 10 is used to obtain the user's current voice information and determine the voice wake-up word according to the current voice information.
[0105] The generation module 20 is used to generate deep object-finding language content according to the current voice information when the voice wake-up word is a object-finding wake-up word.
[0106] The determination module 10 is further configured to obtain a multi-dimensional object feature library of a target area by calling a resource interface, and determine location information of an object corresponding to the deep object search language content based on the multi-dimensional object feature library.
[0107] The search module 30 is used to generate target object-finding guidance information according to the location information of the object, and guide the user to find the object according to the target object-finding guidance information based on the smart speaker.
[0108] This embodiment obtains the user's current voice information and determines a voice wake-up word based on the current voice information; when the voice wake-up word is a find-thing wake-up word, generates deep find-thing language content based on the current voice information; obtains a multi-dimensional item feature library of the target area by calling a resource interface, and determines the location information of the item corresponding to the deep find-thing language content based on the multi-dimensional item feature library; generates target find-thing guidance information based on the location information of the item, and guides the user to find the item based on the target find-thing guidance information based on the smart speaker. Through the above method, after determining the voice wake-up word based on the user's current voice information, it is determined whether the voice wake-up word is a find-thing wake-up word. If so, it indicates that the user needs to find something. At this time, the complex current voice information is deeply analyzed in combination with the context extension information, and then the location information of the item is determined by calling the multi-dimensional item feature library of spatial dimension, video image dimension and radar image dimension. At this time, the smart speaker is used to instruct the user to find the item, thereby effectively improving the accuracy and efficiency of find-things and enhancing the interaction between the user and the smart speaker.
[0109] The intelligent object-finding device provided in this application, which utilizes the intelligent object-finding method described in the aforementioned embodiments, can resolve the technical issues of low accuracy and efficiency in the prior art. Compared to the prior art, the beneficial effects of the intelligent object-finding device provided in this application are the same as those of the intelligent object-finding method described in the aforementioned embodiments, and the other technical features of the intelligent object-finding device are the same as those disclosed in the aforementioned embodiments, and are not further described here.
[0110] In one embodiment, the determination module 10 is further used to obtain the user's current voice information according to a multi-channel sound pickup strategy; perform sound wave recognition on the current voice information to obtain the user's current sound wave information; perform noise reduction on the current sound wave information according to a target noise reduction algorithm, and extract the sound wave feature information of the noise-reduced sound wave information; input the sound wave feature information into a target wake-up word recognition model, and obtain the voice wake-up word recognized and output by the target wake-up word recognition model.
[0111] In one embodiment, the determination module 10 is further configured to generate a target floor plan based on the user's location information if it is detected that the user has not imported the floor plan within a preset time period; perform two-dimensional recognition on the target floor plan, assign three-dimensional parameters to the recognized plane space data, and obtain the current three-dimensional space of the target area; construct a spatial feature database based on the feature information of each item in the current three-dimensional space and the current three-dimensional space; control the camera device to collect video image data of each item in the current three-dimensional space, and determine the first position information of each item based on the item video feature information of the video image data; construct an item video image feature library based on the item video feature information and each first position information; control the radar to collect radar image data of each item in the current three-dimensional space, and determine the second position information of each item based on the item radar feature information of the radar image data; construct an item radar image feature library based on the item radar feature information and each second position information; generate a multi-dimensional item feature library of the target area based on the spatial feature database, the item video image feature library, and the item radar image feature library, and upload the multi-dimensional item features of the target area.
[0112] In one embodiment, the determination module 10 is further used to perform keyword recognition on the deep object search language content to obtain each deep object search keyword; determine the intended item set and predicted probability corresponding to each deep object search keyword based on the item prediction model; obtain the attribute information of the intended items in each of the intended item sets; and determine the location information of the items corresponding to the deep object search language content based on the predicted probability, the attribute information and the multi-dimensional item feature library.
[0113] In one embodiment, the determination module 10 is further used to traverse the same intended items between each of the intended item sets and count the number of the same intended items; when the number is greater than a preset value, filter the items corresponding to the deep object search language content from the same intended items whose number is greater than the preset value based on the predicted probability and the attribute information; and determine the location information of the item based on the multi-dimensional item feature library.
[0114] In one embodiment, the determination module 10 is further used to obtain a spatial item feature information set in the spatial feature database, match the feature information of the item with each feature information in the spatial item feature information set; determine the spatial information where the item is located based on the spatial item feature information matching result; filter the item video image feature library and the item radar image feature library respectively according to the spatial information; obtain the item video feature information set in the filtered item video image feature library, match the feature information of the item with each feature information in the item video feature information; determine the first position information of the item based on the item video feature information matching result; obtain the item radar feature information set in the filtered item radar image feature library, match the feature information of the item with each feature information in the item radar feature information set; determine the second position information of the item based on the item radar feature information matching result, and determine the position information of the item based on the first position information and the second position information.
[0115] In one embodiment, the generation module 20 is also used to perform data cleaning on the current voice information when the voice wake-up word is a search wake-up word; filter the cleaned voice information, and expand the filtered voice information; generate target voice information based on the filtered voice information and the expanded voice information; convert the target voice information to obtain target text information; perform deep reasoning on the target text information according to the natural language processing strategy to obtain deep search language content.
[0116] The present application provides an intelligent object-finding device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the intelligent object-finding method in the above-mentioned embodiment one.
[0117] Reference below Figure 5 , which shows a schematic diagram of the structure of an intelligent object-finding device suitable for implementing the embodiments of the present application. The intelligent object-finding device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5The smart object-finding device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0118] like Figure 5 As shown, the intelligent object-finding device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in ROM (Read Only Memory) 1002 or programs loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the intelligent object-finding device. The processing device 1001, ROM 1002, and RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, hard disk, etc.; and a communication device 1009. Communication device 1009 can allow the intelligent object-finding device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows an intelligent object-finding device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0119] In particular, according to the embodiments disclosed herein, the processes described above with reference to the flowcharts can be implemented as computer software programs. The computer programs contain program code for executing the methods shown in the flowcharts. In such embodiments, the computer programs can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer programs are executed by processing device 1001, the above-described functions defined in the methods of the embodiments disclosed herein are performed.
[0120] The intelligent object-finding device provided in this application, which utilizes the intelligent object-finding method described in the aforementioned embodiment, can resolve the technical issues of low accuracy and efficiency in the prior art. Compared with the prior art, the beneficial effects of the intelligent object-finding device provided in this application are the same as those of the intelligent object-finding method described in the aforementioned embodiment, and the other technical features of the intelligent object-finding device are the same as those disclosed in the method described in the preceding embodiment, and are not further described here.
[0121] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0122] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0123] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. An intelligent object-finding method, characterized in that: The method comprises: Obtaining the user's current voice information and determining a voice wake-up word based on the current voice information; When the voice wake-up word is a find-object wake-up word, generating deep find-object language content according to the current voice information; Acquire a multi-dimensional item feature library of the target area by calling a resource interface, and determine the location information of the item corresponding to the deep object-finding language content based on the multi-dimensional item feature library; Targeted object-finding guidance information is generated according to the location information of the object, and the smart speaker guides the user to find the object according to the targeted object-finding guidance information.
2. The method according to claim 1, wherein The step of obtaining the user's current voice information and determining the voice wake-up word according to the current voice information includes: Acquire the user's current voice information based on the multi-channel voice pickup strategy; Performing sound wave recognition on the current voice information to obtain current sound wave information of the user; Denoising the current sound wave information according to a target noise reduction algorithm, and extracting sound wave feature information of the denoised sound wave information; The acoustic wave feature information is input into a target wake-up word recognition model, and the voice wake-up word recognized and output by the target wake-up word recognition model is obtained.
3. The method according to claim 1, wherein When the voice wake-up word is a find-object wake-up word, the step of generating deep find-object language content according to the current voice information includes: When the voice wake-up word is a find-and-seek wake-up word, performing data cleaning on the current voice information; Filtering the cleaned voice information and performing information expansion on the filtered voice information; generating target voice information according to the filtered voice information and the expanded voice information; Converting the target voice information to obtain target text information; The target text information is subjected to deep reasoning according to a natural language processing strategy to obtain deep object-seeking language content.
4. The method according to any one of claims 1 to 3, characterized in that The step of determining the location information of the object corresponding to the deep object search language content according to the multi-dimensional object feature library includes: Performing keyword recognition on the deep object search language content to obtain deep object search keywords; Determine the intended item set and prediction probability corresponding to each deep object search keyword based on the item prediction model; Obtaining attribute information of each desired item in the desired item set; The location information of the object corresponding to the deep object-finding language content is determined according to the predicted probability, the attribute information and the multi-dimensional object feature library.
5. The method according to claim 4, wherein The step of determining the location information of the object corresponding to the deep object search language content based on the predicted probability, the attribute information, and the multi-dimensional object feature library includes: Traversing the same desired items between each of the desired item sets, and counting the number of the same desired items; When the number is greater than a preset value, selecting items corresponding to the deep object-seeking language content from the same intended items whose number is greater than the preset value according to the predicted probability and the attribute information; The location information of the object is determined according to the multi-dimensional object feature library.
6. The method according to claim 5, wherein The multi-dimensional item feature library includes a spatial feature database, an item video image feature library, and an item radar image feature library; The step of determining the location information of the item based on the multi-dimensional item feature library includes: Obtaining a spatial item feature information set from the spatial feature database, and matching the feature information of the item with each feature information in the spatial item feature information set; Determine the spatial information where the item is located based on the spatial item feature information matching result; filtering the object video image feature library and the object radar image feature library respectively according to the spatial information; Acquire a set of item video feature information from the filtered item video image feature library, and match the feature information of the item with each feature information in the item video feature information; Determining first location information of the object based on a matching result of the object video feature information; Obtaining a set of item radar feature information from the filtered item radar image feature library, and matching the feature information of the item with each feature information in the set of item radar feature information; The second location information of the object is determined according to the matching result of the radar feature information of the object, and the location information of the object is determined according to the first location information and the second location information.
7. The method according to claim 1, wherein Before the step of acquiring the multi-dimensional item feature library of the target area by calling the resource interface, the method further includes: If it is detected that the user has not imported the floor plan within the preset time period, a target floor plan is generated based on the user's location information; Performing two-dimensional recognition on the target floor plan, assigning three-dimensional parameters to the recognized plane space data, and obtaining the current three-dimensional space of the target area; constructing a spatial feature database based on feature information of each object in the current three-dimensional space and the current three-dimensional space; Controlling the camera device to collect video image data of each object in the current three-dimensional space, and determining first position information of each object based on video feature information of the object in the video image data; Building an object video image feature library based on the object video feature information and each of the first position information; Controlling the radar to collect radar image data of each object in the current three-dimensional space, and determining second position information of each object based on radar feature information of the object in the radar image data; Building an object radar image feature library based on the object radar feature information and each of the second position information; A multi-dimensional item feature library of the target area is generated according to the spatial feature database, the item video image feature library and the item radar image feature library, and the multi-dimensional item features of the target area are uploaded.
8. An intelligent object-finding device, characterized in that: The device comprises: A determination module, configured to obtain the user's current voice information and determine a voice wake-up word based on the current voice information; A generation module, configured to generate deep object-finding language content based on the current voice information when the voice wake-up word is a find-object wake-up word; The determination module is further configured to obtain a multi-dimensional item feature library of the target area by calling a resource interface, and determine the location information of the item corresponding to the deep object search language content based on the multi-dimensional item feature library; The search module is used to generate target object-finding guidance information according to the location information of the object, and guide the user to find the object according to the target object-finding guidance information based on the smart speaker.
9. An intelligent object-finding device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the intelligent object-finding method according to any one of claims 1 to 7.
10. A smart speaker, characterized in that: The smart speaker includes: a speaker, an indicator, a display, a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the smart object-finding method as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Searching method, searching device and searching system
CN109029449A
Object searching guide system for special objects
CN111858993A
Positioning method, apparatus and system, and electronic device and storage medium
WO2022100238A1