Intelligent search method, apparatus and intelligent device

By detecting interactive trigger signals in real time and using an edge-cloud collaborative architecture, smart devices have achieved precise positioning and personalized search of target objects in panoramic images, solving the accuracy and real-time issues of existing smart search and improving the user experience.

CN122633893APending Publication Date: 2026-08-25HANGZHOU JIEFENG SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610767354.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing intelligent search solutions suffer from poor target positioning accuracy, weak scene adaptability, high data processing latency, and insufficient user privacy protection, resulting in low search accuracy and real-time performance, and a poor user experience.

Method used

By acquiring panoramic images collected by smart devices, the system can detect interactive trigger signals in real time to locate the target area, extract the feature vector of the target object and add semantic tags, search for related data based on the semantic tags and overlay the results in the panoramic image, improve user intent recognition by adopting a multimodal interaction method, and combine edge-cloud collaborative architecture for lightweight feature extraction and local caching response.

Benefits of technology

It achieves precise location of target objects, improves the accuracy and response speed of intelligent search, ensures user privacy protection, and provides a personalized and scenario-adaptive search experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122633893A_ABST
    Figure CN122633893A_ABST
Patent Text Reader

Abstract

The application provides a kind of intelligent search method, device and intelligent equipment, it is related to the technical field of intelligent search, the method comprises: obtaining panoramic image collected by intelligent device;Real-time detection is triggered for the interactive signal of panoramic image, to locate target area in panoramic image based on interactive trigger signal;Extract the feature vector of target object contained in target area, add semantic label to target object based on feature vector;Search data associated with target object based on semantic label, and superimposed display is carried out in panoramic image with search result.The intelligent search method, device and intelligent equipment provided by the application, the whole search process can be realized by real-time detection triggered for the interactive signal of panoramic image, and the semantic label is added to the target area positioned, the accurate positioning of target object can be realized, and then the accuracy of intelligent search is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of intelligent search, and in particular to an intelligent search method, apparatus, and intelligent device. Background Technology

[0002] With the widespread adoption of smart hardware, visual search has become a key technological direction for overcoming the limitations of traditional text search. However, existing solutions still suffer from numerous technical bottlenecks, including poor target localization accuracy, weak scene adaptability, high data processing latency, and insufficient user privacy protection. These issues result in low accuracy and real-time performance of intelligent search, thus reducing the user experience. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide an intelligent search method, apparatus and intelligent device to alleviate the above-mentioned technical problems and improve the accuracy of search.

[0004] In a first aspect, embodiments of the present invention provide an intelligent search method, the method comprising: acquiring a panoramic image collected by an intelligent device; detecting in real time an interactive trigger signal for the panoramic image, thereby locating a target region in the panoramic image based on the interactive trigger signal; extracting feature vectors of target objects contained in the target region, adding semantic tags to the target objects based on the feature vectors; searching for data associated with the target objects based on the semantic tags, and overlaying the search results in the panoramic image.

[0005] In conjunction with the first aspect, the present invention provides a first possible implementation of the first aspect, wherein the step of acquiring the panoramic image collected by the smart device includes: acquiring the panoramic image collected in real time by the smart device, and displaying the panoramic image collected in real time through an interactive display interface.

[0006] In conjunction with the first aspect, this embodiment of the invention provides a second possible implementation of the first aspect, wherein the step of real-time detection of an interaction trigger signal for the panoramic image to locate a target region in the panoramic image based on the interaction trigger signal includes: detecting the interaction trigger signal in any of the following ways, and locating the target region in the panoramic image based on the interaction trigger signal: real-time detection of a touch operation applied to the panoramic image, generating an interaction trigger signal based on the point of application of the touch operation, and determining the region containing the point of application as the target region of the panoramic image; tracking gaze information applied to the panoramic image, and if gaze information applied to the panoramic image is tracked, determining that the interaction trigger signal has been detected, and determining the target region of the panoramic image based on the point of application of the gaze information; receiving voice information issued for the panoramic image, extracting descriptive information of a target object from the voice information, and if the target object is located in the panoramic image, determining that the interaction trigger signal has been detected, and determining the region containing the target object as the target region.

[0007] In conjunction with the first aspect, this embodiment of the invention provides a third possible implementation of the first aspect, wherein the step of extracting the feature vector of the target object contained in the target region and adding a semantic label to the target object based on the feature vector includes: performing confidence detection on the target region to detect the category confidence of the target object contained in the target region; if the category confidence is greater than a preset confidence threshold, then extracting the feature vector of the target object contained in the target region and adding a semantic label to the target object based on the feature vector.

[0008] In conjunction with the first aspect, or the third possible implementation of the first aspect, this embodiment of the invention provides a fourth possible implementation of the first aspect, wherein the step of extracting the feature vector of the target object contained in the target region and adding semantic tags to the target object based on the feature vector includes: extracting the visual features and semantic features of the target object using a hierarchical feature extraction method, and generating a fusion feature of the target object based on the visual features and the semantic features; and searching for a semantic tag that matches the fusion feature in a pre-set tag library.

[0009] In conjunction with the third possible implementation of the first aspect, this embodiment of the invention provides a fifth possible implementation of the first aspect, wherein the tag library includes a pre-established ternary hierarchical tag library; the step of searching for semantic tags matching the fusion feature in the pre-set tag library includes: searching for multiple semantic tags matching the fusion feature in the ternary hierarchical tag library respectively; and generating a set of semantic tags corresponding to the target object based on the multiple semantic tags.

[0010] In conjunction with the first aspect, this embodiment of the invention provides a sixth possible implementation of the first aspect, wherein the step of searching for data associated with the target object based on the semantic tag includes: searching for historical data matching the semantic tag in a cache library, and determining the historical data with a similarity greater than a similarity threshold to the semantic tag as data associated with the target object; if there is no historical data with a similarity greater than a similarity threshold to the semantic tag in the cache library, then extracting keywords from the semantic tag; and encrypting and sending the keywords to a search engine so that the search engine can search for data associated with the target object.

[0011] In conjunction with the first aspect, this embodiment of the invention provides a seventh possible implementation of the first aspect, wherein the step of overlaying the search results in the panoramic image includes: generating an information display control, adding the search results to the information display control, and overlaying the information display control containing the search results to a preset position in the panoramic image.

[0012] Secondly, embodiments of the present invention also provide an intelligent search device, the device comprising: an acquisition module for acquiring panoramic images collected by an intelligent device; a detection module for detecting interactive trigger signals for the panoramic image in real time, so as to locate a target region in the panoramic image based on the interactive trigger signals; an extraction module for extracting feature vectors of target objects contained in the target region, and adding semantic tags to the target objects based on the feature vectors; and a search module for searching data associated with the target objects based on the semantic tags, and overlaying the search results in the panoramic image.

[0013] Thirdly, embodiments of the present invention also provide an intelligent device, wherein the controller of the intelligent device is configured with the intelligent search device of the second aspect.

[0014] The embodiments of the present invention bring the following beneficial effects: This invention provides an intelligent search method, apparatus, and intelligent device that can acquire panoramic images collected by an intelligent device; detect interactive trigger signals for the panoramic image in real time to locate target areas in the panoramic image based on the interactive trigger signals; extract feature vectors of target objects contained in the target areas and add semantic tags to the target objects based on the feature vectors; search for data associated with the target objects based on the semantic tags, and overlay the search results in the panoramic image. The entire search process can achieve accurate positioning of target objects by detecting interactive trigger signals for the panoramic image in real time and adding semantic tags to the located target areas, thereby improving the accuracy of intelligent search.

[0015] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 A flowchart of an intelligent search method provided in an embodiment of the present invention; Figure 2 A flowchart of another intelligent search method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an intelligent search device provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Traditional intelligent visual search solutions typically suffer from several bottlenecks, including: (1) The accuracy of target positioning is poor.

[0021] Traditional visual search relies primarily on uploading and analyzing entire images, which is prone to recognition errors in multi-target scenarios. For example, when searching for a specific snack on a supermarket shelf, adjacent items might be misidentified. While some search technologies employ interactive tracking, these are limited to head-mounted displays and fail to address the issue of prioritizing multiple targets.

[0022] (2) The scene adaptation capability is relatively weak.

[0023] Most existing search systems are optimized for single scenarios. When faced with complex scenarios such as industrial equipment inspection, the lack of a dedicated feature library leads to low recognition accuracy.

[0024] (3) The data processing latency is high.

[0025] In pure cloud processing mode, the transmission and analysis of images usually takes more than 500ms, which cannot meet the high real-time requirements of industrial inspection and other similar processes.

[0026] (4) There are deficiencies in the protection of user privacy.

[0027] In existing intelligent search processes, raw images are usually uploaded directly to the cloud, which poses a risk of data leakage and lacks optimization mechanisms for personalized search.

[0028] Based on this, the intelligent search method, apparatus, and intelligent device provided in this embodiment of the invention can alleviate the above-mentioned technical problems and improve the accuracy of the search.

[0029] To facilitate understanding of this embodiment, a detailed description of an intelligent search method disclosed in this embodiment of the invention will be provided first.

[0030] In one possible implementation, this invention provides an intelligent search method that can be applied to intelligent devices with camera functions, such as smart terminals (mobile phones, tablets), industrial inspection equipment (inspection robots), smart homes, and other intelligent devices with cameras. This method can meet diverse visual search needs such as product inquiry, equipment fault retrieval, and item information acquisition.

[0031] Specifically, Figure 1 A flowchart of an intelligent search method is shown, which includes the following steps: Step S102: Acquire panoramic images collected by the smart device; In practical use, the smart device in this embodiment of the invention is equipped with an image acquisition unit such as a camera, which can capture panoramic images. For example, taking wearable smart glasses as an example, after the user wears the smart glasses, he / she can turn his / her gaze to a certain direction. At this time, the panoramic view of that direction can be transmitted to the user's field of vision through the camera. Furthermore, the user can input shooting commands, such as manual commands or voice commands, etc., and thus capture the panoramic image of the current field of vision.

[0032] In practical use, to achieve intelligent search, the image acquisition units such as cameras in smart devices are usually high-definition image acquisition units, and are equipped with multi-dimensional environmental perception units and hardware preprocessing units. For example, the high-definition image acquisition unit can use a high-definition CMOS sensor, support a frame rate of 30fps and a focus range of 0.1m-10m, and can integrate macro mode and backlight compensation function; the multi-dimensional environmental perception unit usually includes a light sensor (such as a detection range of 0-65535 lux), a distance sensor (accuracy ±1cm), and a scene light color temperature sensor (2000K-10000K), etc., which can output environmental parameters in real time; the hardware preprocessing unit can be based on FPGA (Field-Programmable Gate Array) chip to implement Gaussian noise reduction, white balance correction, and ROI (Region of Interest) pre-cropping, and can handle latency ≤10ms.

[0033] The specific hardware configuration of the image acquisition unit can be set according to the actual usage, and the embodiments of the present invention do not impose any restrictions on this.

[0034] Step S104: Real-time detection of interactive trigger signals for the panoramic image, so as to locate the target area in the panoramic image based on the interactive trigger signals; Step S106: Extract the feature vector of the target object contained in the target region, and add semantic labels to the target object based on the feature vector; Step S108: Search for data associated with the target object based on semantic tags, and overlay the search results in the panoramic image.

[0035] In this process, step S104 is actually the interaction process between the user and the smart device. For the panoramic image that has been acquired, the user can trigger an interaction signal to locate the target area. The interaction trigger signal can be a signal determined by interaction behavior such as touch, voice, or eye tracking, thereby determining the target area.

[0036] Furthermore, in step S106 above, when adding semantic labels to the feature vector of the target region, a three-layer label library can be called to generate three-layer semantic labels containing name, attribute, and scene association, thereby improving the accuracy of the search.

[0037] Therefore, the intelligent search method provided in this embodiment of the invention can acquire panoramic images collected by intelligent devices; detect interactive trigger signals for the panoramic images in real time, and locate target areas in the panoramic images based on the interactive trigger signals; extract feature vectors of target objects contained in the target areas, and add semantic tags to the target objects based on the feature vectors; search for data associated with the target objects based on the semantic tags, and overlay the search results in the panoramic images. The entire search process can achieve accurate positioning of target objects by detecting interactive trigger signals for the panoramic images in real time and adding semantic tags to the located target areas, thereby improving the accuracy of intelligent search.

[0038] In practical use, the process of acquiring panoramic images in step S102 above is actually a process of real-time acquisition of panoramic images. That is, panoramic images are acquired in real time through smart devices and displayed through an interactive display interface. Again, taking a user wearing smart glasses as an example, suppose a user is wearing interactive smart glasses in a supermarket setting. The smart glasses have an interactive display interface set up in the user's glasses area, and a camera is set up on the outward-facing side of the smart glasses to capture a panoramic view of the shelves. For example, if a user wants to search for information about a beverage in the beverage section, such as its ingredients or nutritional components, the camera's field of view can be pointed at the beverage shelves to capture a real-time view of the beverage section. The user can then see this view on the interactive display interface of the smart glasses. When the user sees the beverage they want to search for, they can interact with the smart glasses, allowing the smart glasses to recognize the user's intent.

[0039] In practical use, the process by which smart devices recognize user intent and add semantic tags to target objects is actually a multimodal processing procedure. Users can interact with smart devices in various ways, enabling the devices to recognize user intent and obtain the target region. This allows for further feature extraction to obtain feature vectors from the target region. Therefore, the main task of this multimodal processing procedure is to realize the recognition of user intent and the extraction of visual features. For ease of understanding, in Figure 1 On this basis, Figure 2 A flowchart of another intelligent search method is shown, further describing the multimodal processing procedure, specifically, as follows: Figure 2 As shown, it includes the following steps: Step S202: Acquire panoramic images collected in real time by the smart device, and display the real-time panoramic images through an interactive display interface. During this process, a multi-dimensional environmental perception unit can be activated simultaneously. For example, the multi-dimensional environmental perception unit can simultaneously acquire illumination values, distance values, and color temperature values. If the illumination value is <50 lux, a supplementary light (brightness adjustable from 10-100 lm) is turned on; if the color temperature deviation is >1000K, white balance correction is triggered, etc., so that the panoramic image is more suitable for subsequent processes such as intent recognition. The specific adjustment process of the environmental perception unit can be set according to the actual use situation, and the embodiments of the present invention do not limit this.

[0040] Furthermore, in this embodiment of the invention, for the acquired panoramic image, the interactive trigger signal can be detected by any of the following methods, and the target area in the panoramic image can be located based on the interactive trigger signal, specifically including steps S204 to S208.

[0041] Step S204: Real-time detection of touch operations applied to the panoramic image; generation of interactive trigger signals based on the point of action of the touch operation; and determination of the area containing the point of action as the target area of ​​the panoramic image. Step S206: Track the gaze information acting on the panoramic image. If gaze information acting on the panoramic image is tracked, it is determined that an interaction trigger signal has been detected, and the target area of ​​the panoramic image is determined based on the point of action of the gaze information. Step S208: Receive voice information sent to the panoramic image, extract description information of the target object from the voice information, and if the target object is located in the panoramic image, determine that an interaction trigger signal has been detected and define the area containing the target object as the target area. In practical use, to achieve the interactive process described in steps 204 to S208, smart devices are typically equipped with corresponding interactive algorithms. For example, a multimodal target localization algorithm can be set to recognize touch operations, including recognizing gesture boxes and fingertip coordinates to determine the point of action, and generating bounding boxes based on the target objects corresponding to the fingertip coordinates to determine the target area, etc. In addition, a gaze estimation algorithm can be set, such as the Gaze point localization algorithm, which can determine the position of the user's gaze point in the panoramic image through eye tracking, and can also determine the target area. Furthermore, a speech recognition algorithm can be set, that is, to recognize the user's speech information and extract descriptive information, such as "yellow bottle", "third beverage bottle", etc., and then locate the target object in the panoramic image based on the descriptive information.

[0042] Furthermore, for the detection methods in steps S204 to S208, they can be enabled simultaneously or at least one can be selected for use. The specific settings can be configured according to actual usage, and this embodiment of the invention does not impose any limitations on this. Moreover, when different detection methods are used simultaneously, if any one method detects and identifies the target area, subsequent steps can be executed. Furthermore, the detection results from various methods can be uniformly encoded into structured vectors. For example, when encoding gesture intent, not only can the target region be identified, but also the attention weight and operation type (op_type, such as "search" or "compare") encoded by gesture trajectory (e.g., drawing a circle for emphasis) and hand shape (e.g., pointing with the index finger) can be included. For eye-tracking, the user's confidence level and focus area (e.g., product logo or device nameplate) can be encoded by the user's gaze duration and pupil changes. For speech information, descriptive keywords (desc_keywords) and action commands (action_cmd), such as "What model is this?" or "Is it broken?", can be extracted. All of the above information can then be uniformly encoded into an intent vector for output, facilitating subsequent feature extraction.

[0043] In addition, the target area can be preprocessed by the hardware preprocessing unit, such as reducing noise in the target area and cropping the target area into a standard-sized image, etc. The specific settings can be made according to the actual use situation, and the embodiments of the present invention do not limit this.

[0044] Furthermore, the feature process in the embodiments of the present invention is actually a secondary detection process of the target region, specifically including the following steps.

[0045] Step S210: Perform confidence detection on the target region to detect the category confidence of the target objects contained in the target region; Step S212: If the category confidence is greater than the preset confidence threshold, extract the feature vector of the target object contained in the target region, and add semantic labels to the target object based on the feature vector; In practical use, the above steps S210~S212 actually include the process of feature extraction and adding semantic labels. In this embodiment of the invention, when performing feature extraction, an improved YOLOv8 algorithm can be used to detect small targets and output the corresponding category confidence. When the category confidence is greater than the preset confidence threshold, it means that the recognition of the target object is relatively accurate and further feature extraction can be performed. If the category confidence is lower than the confidence threshold, it means that the recognized target object is not accurate and the subsequent search process is not performed. The user can re-determine the target object or add further descriptive information to identify target objects with higher category confidence.

[0046] Furthermore, in this embodiment of the invention, a hierarchical feature extraction method is adopted, including extracting low-level visual features (such as texture and color) and high-level semantic features (such as shape and function), and then performing feature normalization processing to generate the feature vector in this embodiment of the invention.

[0047] That is, the visual and semantic features of the target object are extracted using a hierarchical feature extraction method, and the fused features of the target object are generated based on the visual and semantic features; then, semantic tags that match the fused features are searched in a pre-set tag library.

[0048] Furthermore, the tag library in this embodiment of the invention includes a pre-established ternary hierarchical tag library; when searching for semantic tags that match the fusion features, multiple semantic tags that match the fusion features can be searched in the ternary hierarchical tag library respectively; based on multiple semantic tags, a set of semantic tags corresponding to the target object is generated.

[0049] Specifically, the three-tiered tag library in this embodiment of the invention uses a general library-domain library-user library mechanism to add tags to target objects. The general library provides basic tags for the target object, the domain library adds domain-specific corpus to the target object, and the user library provides personalized tags based on user preferences or personalized information set by the user. Taking a beverage search as an example, the general library first provides basic tags, such as "bottle," "beverage," and "tea"; the domain library can load domain-specific corpus to provide more professional tags, such as "SKU: 6901234567890," "Category: Sugar-free tea beverage," and "Packaging: PET plastic bottle"; the user library can remember previously set or selected user preferences, such as "low-sugar" or "healthy" products, or obtain user-set tag information to generate personalized tags, such as "Attributes you may be interested in: 0 sugar 0 fat," "Review," etc.

[0050] Finally, all the tags provided by the three-element hierarchical tag library are merged to generate a comprehensive semantic tag set, such as "{Name: XX brand oolong tea}, {Attributes: sugar-free, 0 calories, PET bottle}, {Scenario: supermarket, beverage section}, {Extension: user-frequently viewed reviews}", etc. This semantic tag set can guide users in subsequent search steps.

[0051] Step S214: Search for data associated with the target object based on semantic tags; Specifically, the search process in this embodiment of the invention includes local search and cloud search. Local search refers to a local cache library; that is, historical data matching semantic tags is first searched in the cache library, and historical data with a similarity greater than a similarity threshold is identified as data associated with the target object. If no historical data with a similarity greater than the similarity threshold is found in the cache library, keywords are extracted from the semantic tags. The keywords are then encrypted and sent to a search engine so that the search engine can search for data associated with the target object. This search engine is the engine used for cloud search.

[0052] Specifically, the search engine in this embodiment of the invention can employ a dynamic attention alignment mechanism and a scene-aware semantic routing and generation mechanism. The dynamic attention alignment mechanism uses a dynamic attention alignment algorithm to address the difficulty of decoupling and aligning multi-granularity features. For example, during the search process, keywords from the semantic tag set, as well as visual and semantic features from the fusion features, can be extracted. The visual and semantic features are projected into a high-dimensional semantic space and aligned using a cross-attention mechanism to obtain visual features and rich intent information. The semantic routing and generation mechanism further classifies the visual features and rich intent information output by the dynamic attention alignment mechanism according to a preset scene classifier, determining which scene it belongs to. Then, based on the semantic tags of the ternary hierarchical tag library, intelligent routing logic is used for intelligent search. For example, for general layer retrieval of a general library: retrieval is performed in a general knowledge base (such as ImageNet concepts) to obtain the basic category tag L_general; for domain layer routing retrieval of a domain library: based on the scene classification results, dynamic activation and routing to the corresponding domain knowledge base are performed. For example, after identifying an "industrial scenario," the search can prioritize the "part model-fault feature" database rather than the "product SKU" database to obtain the domain-specific tag L_domain, such as "bearing model: 6204-2RS, potential fault: insufficient lubrication." For personalized weighted search of the user layer in the user database, the general and domain tags can be personalized by combining the user's historical profile to obtain search results L_user that meet the user's preferences. Then, semantic generation is performed on all search results. At this point, a lightweight semantic generation network (such as a small Transformer) can be used to generate the final search results by using fused features as conditions and the above L_general, L_domain, and L_user as context. The search results are usually structured semantic tags S and contain rich descriptions of object identity, attributes, and contextual status / suggestions, such as: "This is an Apple iPhone 15 (identity), dark blue (attribute), screen has a crack (status), nearby repair shops (contextual suggestion)", etc.

[0053] In practical use, the search process in this embodiment of the invention is actually a process of intelligent data interaction. Furthermore, to achieve data interaction, intelligent devices can typically be configured with a multi-protocol communication unit, an encryption unit, and an active caching unit. The multi-protocol communication unit can support short-range communication such as WiFi, as well as long-range outdoor and wired communication such as 5G NR and Ethernet transmission. In addition, it can employ TCP / IP protocol encapsulation to ensure the reliability and orderliness of data transmission.

[0054] Furthermore, the encryption unit encrypts the data when sending keywords to the search engine. For example, the SM4 algorithm can be used to encrypt the keywords, and the key is stored in a preset hardware security module (HSM). The encrypted data is then sent to the cloud-based search engine via the communication method of the aforementioned multi-protocol communication unit. In this case, the smart device can be regarded as an edge device, enabling edge-cloud interaction with the cloud-based search engine. In addition, an update cycle can be set for the key to ensure the security of data transmission and storage by updating the key periodically.

[0055] Furthermore, during edge-cloud interaction, user authentication can be performed on the smart device side. For example, the smart device, acting as an edge-end search engine, performs two-way authentication using algorithms such as SM2 before communicating with the cloud. Only after successful authentication is data searching allowed, preventing unauthorized access by unauthorized users. Specifically, the authentication process can be triggered when the smart device starts up or attempts to connect to the cloud. At this time, the smart device can send an identity request signed with its private key (SM2 algorithm) to the cloud. Upon receiving the request, the cloud verifies the signature using the edge's public key, confirming that it is a legitimate and authorized smart device. Then, the cloud also sends its own identity certificate to the edge-end smart device. After the edge verifies the cloud's certificate, two-way authentication can be achieved. At this point, a secure and encrypted communication channel can be established. All data transmitted thereafter, such as encrypted keywords sent to the search engine and search results returned by the search engine, will be encrypted using the SM4 symmetric key negotiated during the authentication process.

[0056] Further proactive caching units can expand the local cache library. Typically, proactive caching units can optimize caching strategies based on multi-scale feature fusion prediction mechanisms, reducing cloud dependency. They are usually designed as follows: Feature extraction is performed on three types of features collected from historical search data: time features (search period, weekly frequency, monthly frequency), scene features (ambient lighting, scene type), and user features (age, gender, search preferences). An encoder-decoder system using an LSTM (Long Short-Term Memory) network is employed for popularity prediction; for example, by inputting multi-scale fusion features, it predicts high-frequency search content for the next 7 days. During caching, the feature vectors and semantic tags of the predicted high-frequency content, along with the search results, are pre-cached to an edge-based cache database. When the cache is full, data with low popularity and no access is prioritized for eviction; for example, data with a predicted popularity <10% and no access for 3 days is prioritized for eviction. Based on this proactive caching unit, a high local cache hit rate of ≥85% can be achieved, with nearly 60% of search requests directly accessing the cache database to obtain search results, resulting in lower response latency, which can be reduced to within 80ms, thereby improving search efficiency.

[0057] Step S216: Generate an information display control, add the search results to the information display control, and overlay the information display control containing the search results onto a preset position of the panoramic image.

[0058] In practical use, the aforementioned smart devices can achieve adaptive interaction through an interactive display interface. This interface typically functions as a human-computer interaction interface, enabling multiple triggering methods and intuitive result presentation. For example, in steps S204-S208, for gesture interaction, the MediaPipe gesture recognition algorithm can be used to identify the point of application of the touch operation, supporting operations such as "box selection" (thumb and forefinger forming a rectangle) and "confirm" (fist clenching), with a trigger response time of ≤20ms. For eye-tracking algorithms, an infrared eye-tracking algorithm can be integrated, automatically generating a target area bounding box after the user gazes at the target for 2 seconds, suitable for scenarios where both hands are busy (such as industrial operations). For voice information triggering, wake words such as "search this" and "identify target" can be supported, combined with voice positioning commands (such as "top left corner item") to improve recognition accuracy. The above interaction process can be manually switched by the user, or priority settings can be supported, such as default gesture > eye tracking > voice information recognition, or all three methods can be used simultaneously, depending on the actual usage. This embodiment of the invention does not impose any limitations on this.

[0059] Furthermore, in step S216, when receiving search results, AR augmented display technology can be used, namely optical perspective AR technology or screen AR overlay technology, to achieve an immersive presentation of search results. For example, an information display control can be generated, which can be an overlay window or a virtual window. The search results can be overlaid on the target object in the form of semi-transparent virtual labels. The content displayed in the information display space can include the target name, core attributes (such as "Price: 99 yuan", "Model: M201") and operation entry points ("View details", "Add to favorites"), etc. Users can interact with the information display control, which can support touch label switching details page, two-finger zoom to adjust label size, and drag label position to avoid obscuring the target, etc. When there are multiple search results, the labels can be arranged in descending order of confidence, and the label of the target object with the highest confidence is highlighted.

[0060] Furthermore, a search relevance control can be added to the information display control. For example, two buttons, "Relevant" and "Irrelevant," can be placed in a preset position at the bottom of the information display control. Users can provide feedback on whether the search results contain the information they need through real-time feedback via the search relevance control. The search relevance control also supports input, allowing users to add supplementary descriptive text, such as "Not this style, I need one with a hat," to optimize the search results. User feedback can be automatically recorded. With user authorization, clicks, favorites, and shares can also be recorded as implicit feedback data, which can expand the local cache.

[0061] In summary, the intelligent search method provided by this invention can accurately capture user intent through multimodal interactions (such as gestures, gaze, and voice), construct fused features including visual and semantic features, and achieve accurate semantic understanding and adaptive matching across scenarios. Employing an edge-cloud collaborative architecture, edge-side intelligent devices are responsible for real-time intent parsing, lightweight feature extraction, and local caching, while the cloud-based search engine is responsible for deep semantic mapping, cross-scenario knowledge fusion, and federated learning optimization. This effectively addresses the problems of ambiguous intent, poor scenario adaptability, and large semantic gaps in traditional visual search. While ensuring high response speed and accuracy, it achieves true scenario adaptation and personalized search, not only enabling precise location of target objects but also improving the accuracy of intelligent search, thereby enhancing the user experience.

[0062] Furthermore, based on the above embodiments, this invention also provides an intelligent search device, such as... Figure 3 The diagram shows the structure of an intelligent search device, which includes: The acquisition module 30 is used to acquire panoramic images collected by the smart device; Detection module 32 is used to detect interactive trigger signals for the panoramic image in real time, so as to locate the target area in the panoramic image based on the interactive trigger signals; Extraction module 34 is used to extract the feature vector of the target object contained in the target region, and add semantic tags to the target object based on the feature vector; Search module 36 is used to search for data associated with the target object based on the semantic tags, and to overlay the search results in the panoramic image.

[0063] Furthermore, this embodiment of the invention also provides an intelligent device, the controller of which is configured with the aforementioned intelligent search device.

[0064] The apparatus and intelligent device provided in the embodiments of the present invention have the same technical features as the intelligent search method provided in the above embodiments, so they can also solve the same technical problems and achieve the same technical effects.

[0065] Furthermore, embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0066] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described method.

[0067] Furthermore, embodiments of the present invention also provide a schematic diagram of the structure of an electronic device, such as... Figure 4 The diagram shows the structure of the electronic device, which includes a processor 41 and a memory 40. The memory 40 stores computer-executable instructions that can be executed by the processor 41, and the processor 41 executes the computer-executable instructions to implement the above-described method.

[0068] exist Figure 4 In the illustrated embodiment, the electronic device further includes a bus 42 and a communication interface 43, wherein the processor 41, the communication interface 43, and the memory 40 are connected via the bus 42.

[0069] The memory 40 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 43 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 42 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 42 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0070] Processor 41 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 41 or by instructions in software form. Processor 41 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this invention can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory, and the processor 41 reads the information in the memory and uses its hardware to complete the aforementioned method.

[0071] The computer program product of the intelligent search method, apparatus and intelligent device provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.

[0072] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the above-described device and smart equipment can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0073] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0074] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0075] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0076] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An intelligent search method, characterized in that, The method includes: Acquire panoramic images captured by smart devices; Real-time detection of interactive trigger signals for the panoramic image, so as to locate the target area in the panoramic image based on the interactive trigger signals; Extract the feature vectors of the target objects contained in the target region, and add semantic labels to the target objects based on the feature vectors; Data associated with the target object is searched based on the semantic tags, and the search results are overlaid and displayed in the panoramic image.

2. The method according to claim 1, characterized in that, The steps for acquiring panoramic images captured by smart devices include: The system acquires panoramic images collected in real time by the intelligent device and displays these images through an interactive display interface.

3. The method according to claim 1, characterized in that, The step of real-time detection of interactive trigger signals for the panoramic image, and locating a target region in the panoramic image based on the interactive trigger signals, includes: The interaction trigger signal is detected using any of the following methods, and the target region in the panoramic image is located based on the interaction trigger signal: Real-time detection of touch operations applied to the panoramic image; generation of interactive trigger signals based on the point of application of the touch operations; and determination of the region containing the point of application as the target region of the panoramic image. Track the gaze information acting on the panoramic image. If the gaze information acting on the panoramic image is tracked, it is determined that the interaction trigger signal has been detected, and the target area of ​​the panoramic image is determined based on the point of action of the gaze information. The system receives voice information directed at the panoramic image, extracts description information of the target object from the voice information, and if the target object is located in the panoramic image, it determines that the interaction trigger signal has been detected and identifies the region containing the target object as the target region.

4. The method according to claim 1, characterized in that, The step of extracting feature vectors of target objects contained in the target region and adding semantic labels to the target objects based on the feature vectors includes: A confidence test is performed on the target region to detect the category confidence of the target object contained in the target region; If the category confidence score is greater than a preset confidence score threshold, then the feature vector of the target object contained in the target region is extracted, and a semantic label is added to the target object based on the feature vector.

5. The method according to claim 1 or 4, characterized in that, The step of extracting feature vectors of target objects contained in the target region and adding semantic labels to the target objects based on the feature vectors includes: The visual and semantic features of the target object are extracted using a hierarchical feature extraction method, and a fusion feature of the target object is generated based on the visual and semantic features. Search for semantic tags that match the fusion features in a pre-set tag library.

6. The method according to claim 5, characterized in that, The tag library includes a pre-established ternary hierarchical tag library; The step of searching for semantic tags that match the fused features in a pre-set tag library includes: Search for multiple semantic tags that match the fusion feature in the ternary hierarchical tag library; Based on the multiple semantic tags, a set of semantic tags corresponding to the target object is generated.

7. The method according to claim 1, characterized in that, The step of searching for data associated with the target object based on the semantic tags includes: Search the cache library for historical data that matches the semantic tag, and determine the historical data that has a similarity greater than the similarity threshold with the semantic tag as the data associated with the target object; If there is no historical data in the cache library that has a similarity greater than the similarity threshold with the semantic tag, then keywords are extracted from the semantic tag; The keywords are encrypted and sent to a search engine so that the search engine can search for data associated with the target object.

8. The method according to claim 1, characterized in that, The step of overlaying the search results onto the panoramic image includes: Generate an information display control, add the search results to the information display control, and overlay the information display control containing the search results onto a preset position of the panoramic image.

9. An intelligent search device, characterized in that, The device includes: The acquisition module is used to acquire panoramic images collected by smart devices; The detection module is used to detect interactive trigger signals for the panoramic image in real time, so as to locate the target area in the panoramic image based on the interactive trigger signals; The extraction module is used to extract the feature vectors of the target objects contained in the target region, and add semantic tags to the target objects based on the feature vectors; The search module is used to search for data associated with the target object based on the semantic tags, and to overlay the search results in the panoramic image.

10. A smart device, characterized in that, The controller of the smart device is equipped with the smart search device as described in claim 9.