Target searching and tracking method and device, storage medium and electronic equipment

By searching for target features in the target image and using feature search models for feature extraction and positioning, the problem of high computing power requirements in the existing technology is solved, and efficient and accurate target tracking is achieved under low computing power.

CN120236098APending Publication Date: 2025-07-01SHENZHEN SANLIJIE INTELLIGENT MANUFACTURING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510285386.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing target tracking technologies require a large amount of computing resources, resulting in high demand for computing power of electronic devices and inability to run other applications.

Method used

Search the target features in the target sub-graph to obtain feature information, and use the feature search model to extract and locate feature to reduce computing power requirements.

Benefits of technology

It realizes target tracking with less computing power, is suitable for more applications, improves computing efficiency and accuracy, and adapts to diverse scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236098A_ABST
    Figure CN120236098A_ABST
Patent Text Reader

Abstract

The invention provides a target searching and tracking method and device, a storage medium and electronic equipment, and the method is applied to the technical field of computers, and comprises the steps: obtaining a target sub-graph input by a user, carrying out the feature extraction processing of the target sub-graph to obtain target feature information, obtaining a target image input by the user, and carrying out the feature extraction processing of the target sub-graph; and searching the target feature information in the target image, and if the target feature information exists in the target image, obtaining position information of the target feature information in the target image. According to the method, the position of the target feature can be determined by searching in the target image through the target feature in the target sub-image, the feature information is extracted so as to achieve the purpose of tracking the target feature with small computing power, the method is suitable for more application programs, and the computing power requirement for electronic equipment is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more particularly, to a target search and tracking method, apparatus, storage medium, and electronic device in the field of computer technology. Background Art

[0002] Target tracking technology plays an important role in fields such as security monitoring, intelligent transportation, or human-computer interaction. It can detect and locate targets using image or video data. However, the methods for implementing target tracking technology in the prior art often require a large amount of computing resources to support their operation, which makes them only able to run on electronic devices with high computing power and will occupy a large amount of operating memory of the electronic device, making it impossible for the electronic device to run other applications. There is a need to provide a target tracking method with less computing power and smaller size. Summary of the Invention

[0003] Embodiments of this application provide a target search and tracking method, apparatus, storage medium, and electronic device. This method can search in the target image through the target features in the target subgraph to determine the position of the target features, extract feature information to achieve the purpose of tracking the target features with less computing power, is applicable to more applications, and reduces the computing power requirements for the electronic device.

[0004] In a first aspect, embodiments of this application provide a target search and tracking method, and the method includes: Obtain the target subgraph input by the user, and perform feature extraction processing on the target subgraph to obtain target feature information; Obtain the target image input by the user, and perform search processing on the target feature information in the target image; If the target feature information exists in the target image, obtain the position information of the target feature information in the target image.

[0005] Through the above technical solution, search in the target image through the target features in the target subgraph to determine the position of the target features, extract feature information to achieve the purpose of tracking the target features with less computing power, is applicable to more applications, and reduces the computing power requirements for the electronic device.

[0006] In combination with the first aspect, in some possible implementation manners, before obtaining the target subgraph input by the user and performing feature extraction processing on the target subgraph to obtain target feature information, it further includes: Construct an initial feature search model, obtain a sample subgraph, a sample image, and the sample position information of the sample feature information of the sample subgraph in the sample image; Input the sample sub - graph and the sample image into the initial feature search model to obtain the training position information output by the initial feature search model; Based on the training position information and the sample position information, train the initial feature search model until the initial feature search model meets the training termination condition to obtain the feature search model.

[0007] Through the above technical solution, the initial feature search model is trained specifically through the sample position information, and the model can stop training in time after reaching the termination condition, avoiding ineffective training iterations, saving training time, and improving training efficiency.

[0008] Combined with the first aspect, in some possible implementation manners, the obtaining the target sub - graph input by the user and performing feature extraction processing on the target sub - graph to obtain target feature information includes: Obtain the target sub - graph input by the user; Input the target sub - graph into the target sub - graph branch of the feature search model, and perform feature extraction processing on the target sub - graph to obtain target feature information; Store the target feature information in the feature database of the feature search model.

[0009] Through the above technical solution, the target sub - graph branch is used to specifically perform feature extraction processing on the target sub - graph, improving the quality and efficiency of target feature information extraction. And extracting and storing the target feature information input by the user can meet the personalized needs of the user and improve the adaptability of the model in different scenarios.

[0010] Combined with the first aspect, in some possible implementation manners, the obtaining the target image input by the user and performing search processing on the target feature information in the target image includes: Obtain the target image input by the user, and obtain the target feature information from the feature database; Input the target image into the search network of the feature search model, perform associated convolution processing on the target image and the target feature information, and determine whether the target feature information exists in the target image.

[0011] Combined with the first aspect, in some possible implementation manners, the inputting the target image into the search network of the feature search model and performing associated convolution processing on the target image and the target feature information includes: Input the target image into the search network of the feature search model; Control the search network to perform a convolution operation on the target image using the target feature information as a convolution kernel, and obtain associated feature information that matches the target feature information.

[0012] Through the above technical solution, accurate matching and rapid positioning of the target feature information can be achieved through associated convolution processing, reducing missed detection cases and improving search efficiency.

[0013] Combined with the first aspect, in some possible implementation manners, if the target feature information exists in the target image, obtaining the position information of the target feature information in the target image includes: If there is associated feature information that matches the target feature information, determine that the target feature information exists in the target image, and obtain the category information of the target feature information; Obtain the position information of the associated feature information in the target image; Output the category information and the position information.

[0014] Through the above technical solution, obtaining the category information of the target feature information helps to improve the accuracy of target search and tracking, avoid confusing targets of different categories, improve the reliability of search results, and enable the target search and tracking device to implement search and tracking for multiple targets, making it more adaptable to diverse application scenarios.

[0015] Combined with the first aspect, in some possible implementation manners, training the initial feature search model based on the training position information and the sample position information until the initial feature search model meets the training termination condition to obtain the feature search model includes: Train the initial feature search model based on the training position information and the sample position information until the initial feature search model meets the overfitting termination condition, and output parameter adjustment prompt information for the initial feature search model.

[0016] Through the above technical solution, combined with the overfitting termination condition, the model is prevented from overlearning the noise and details in the sample data, thereby maintaining good generalization ability.

[0017] In a second aspect, an embodiment of the present application provides a target search and tracking device, and the device includes: A target feature acquisition unit, configured to acquire a target subgraph input by a user, and perform feature extraction processing on the target subgraph to obtain target feature information; A target feature search unit, configured to acquire the target image input by the user, and perform search processing on the target feature information in the target image; A location information determination unit, configured to, if the target feature information exists in the target image, obtain the location information of the target feature information in the target image.

[0018] In a third aspect, an embodiment of the present application provides a computer storage medium storing multiple instructions adapted to be loaded and executed by a processor to perform the above method steps.

[0019] In a fourth aspect, an embodiment of the present application provides an electronic device, which may include: a processor and a memory; wherein, the memory stores a computer program adapted to be loaded and executed by the processor to perform the above method steps.

[0020] In one or more embodiments of the present application, a target sub-graph input by a user is obtained, target feature information is obtained by performing feature extraction processing on the target sub-graph, a target image input by the user is obtained, and the target feature information is searched for in the target image. If the target feature information exists in the target image, the location information of the target feature information in the target image is obtained. Searching for the location of the target feature in the target image through the target feature in the target sub-graph, and extracting feature information to achieve the purpose of tracking the target feature with less computing power, which is applicable to more application programs and reduces the computing power requirements for the electronic device. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic structural diagram of a feature search model provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of a target search and tracking method provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of feature search model training provided by an embodiment of the present application; Figure 4 It is a schematic flowchart of feature extraction processing provided by an embodiment of the present application; Figure 5 It is a schematic flowchart of search processing provided by an embodiment of the present application; Figure 6 It is a schematic flowchart of location information acquisition provided by an embodiment of the present application; Figure 7 It is a schematic structural diagram of a target search and tracking device provided by an embodiment of the present application; Figure 8 It is a schematic structural diagram of a target search and tracking device provided by an embodiment of the present application; Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0024] The target search and tracking method provided by the embodiment of the present application can be implemented depending on a computer program, and can run on a target search and tracking device based on the von Neumann architecture. This computer program can be integrated in an application or run as an independent tool application. The target search and tracking device provided by the embodiment of the present application can be installed on electronic devices such as mobile phones, smart glasses, and cameras, and can achieve the purpose of searching and tracking the target by occupying less computing resources and running memory. For example, the user can input a target sub-image into the target search and tracking device, and the target search and tracking device can obtain the target feature information that the user wants to track from the target sub-image. The target sub-image is an image in which there is a target that the user wants to track, and the target feature information is the key feature that can accurately describe the target, and is the feature information that can effectively distinguish the target in a complex background or scene. The target in the target sub-image can be a human body, a human face, an animal, or an object, etc. The target sub-image can be determined by the user according to the requirements, according to the type of the electronic device, and according to the current usage scenario. After that, the user can input a target image into the target search and tracking device, and the target search and tracking device can determine the position information of the target feature information in the target image. Among them, the position information can include the coordinates of the target feature information in the target image, and the position information can also be a bounding box, and the target feature information is in the bounding box.

[0025] In order to improve the calculation efficiency and accuracy of target search and tracking, the target search and tracking device can adopt a feature search model to obtain the position information. Please also refer to Figure 1, This application's embodiment provides a structural schematic diagram of a feature search model. The feature search model may include a backbone network, a search network, a target sub-graph branch, and a feature database. Among them, the backbone network is used to extract features from pictures and can be composed of several components with different depths, convolutions, and sizes. Each component may include a convolutional layer, a pooling layer, a BN layer, and an activation function. The search network can be used to obtain the location information of target feature information in the target image and can be composed of two layers of components. The target sub-graph branch can be used to extract the target feature information in the target sub-graph. After the user inputs the target sub-graph and the target image into the target search and tracking device, the target search and tracking device can input both the target sub-graph and the target image into the feature search model. The backbone network can perform preliminary feature extraction on the target sub-graph to obtain sub-graph features, and then input the sub-graph features into the target sub-graph branch. The target sub-graph branch can extract the refined features of the target sub-graph based on the sub-graph features, thereby obtaining the target feature information in the target sub-graph and inputting it into the feature database for storage. The backbone network can input the target image into the search network. The search network can obtain the target feature information from the feature database and then obtain the location information of the target feature information in the target image. The target search and tracking device can obtain the location information output by the search network.

[0026] Since the feature database can store one target feature information, it can store multiple target feature information, and different target feature information can come from different target sub-graphs. Therefore, the search network can output the location information and category information of each target feature information. The category information is used to indicate whether there is target feature information in the target image. The category information can be set by the target search and tracking device or can be set and named by the user. For example, the category information can be 0 or 1. 0 indicates that there is no such target feature information in the target image, and 1 indicates that there is target feature information in the target image. The location information indicates the location of the target feature information in the target image. It can be understood that there will only be the location information corresponding to the target feature information when the category information is 1. For example, if there is target feature information A and target feature information B in the data feature library, the search network can sequentially output the location information and category information of each target feature information corresponding to the target image. If there is target feature information A in the target image, it can output category information A and location information A, where category information A can be 1 and location information A is the location of target feature information A in the target image. If there is no target feature information B in the target image, it can output category information B, where category information B can be 0. Since there is no target feature information B, there is no corresponding location information.

[0027] The following will specifically describe the target search and tracking method provided by this application with specific embodiments.

[0028] Please refer to Figure 2, which provides a schematic flowchart of a target search and tracking method for an embodiment of the present application. As Figure 2 shown, the method of the embodiment of the present application may include the following steps S101 - S103.

[0029] S101, obtain the target sub - graph input by the user, and perform feature extraction processing on the target sub - graph to obtain target feature information.

[0030] Specifically, the user can input the target sub - graph containing the target to be tracked into the target search and tracking device. After the target search and tracking device obtains the target sub - graph input by the user, it can perform feature extraction processing on the target sub - graph to obtain target feature information. Among them, the target sub - graph can be one or more pictures, and each target sub - graph can contain one target feature information. The target search and tracking device can search for all target feature information in the target image.

[0031] S102, obtain the target image input by the user, and perform search processing on the target feature information in the target image.

[0032] Specifically, the target search and tracking device can obtain the target image input by the user and perform search processing on the target feature information in the target image. For example, the target search and tracking device can extract the image feature information in the target image, and then match the target feature information with the image feature information. If there is image feature information that matches the target feature information, it means that there is target feature information in the target image. The target image can be one or more images. For example, the user can input multiple target images, and the target search and tracking device can perform search processing on all target images. The user can also input video data, and the target image is each frame image in the video data, so as to realize the tracking of the target in the video.

[0033] S103, if there is target feature information in the target image, obtain the position information of the target feature information in the target image.

[0034] Specifically, if the target search and tracking device detects the target feature information in the target image, that is, there is target feature information in the target image, it can obtain the position information of the target feature information in the target image. For example, if there is image feature information that matches the target feature information in the target image, the position information of the target sub - graph can be regressed from the target feature information, that is, the position information of this image feature information in the target image is the position information of the target feature information in the target image.

[0035] Among them, the position information can be the coordinate information in the target image regressed from the target feature information, or the position information of the annotation box where the target feature information is located. For example, the position information can include the upper left vertex coordinates of the annotation box, the length of the annotation box, and the width of the annotation box. The target search and tracking device can generate an annotation box in the target image according to the position information and display the target image with the annotation box to the user, so as to intuitively display the target in the target image to the user and achieve the search and tracking of the target.

[0036] In the embodiment of the present application, the target sub-graph input by the user is obtained, the target feature information is obtained by performing feature extraction processing on the target sub-graph, the target image input by the user is obtained, and the target feature information is searched in the target image. If the target feature information exists in the target image, the position information of the target feature information in the target image is obtained. Searching for the position of the feature of the target sub-graph in the target image through the target feature in the target sub-graph, extracting feature information to achieve the purpose of tracking a specific target with less computing power, being applicable to more applications, and reducing the computing power requirements for the electronic device.

[0037] The target search and tracking device can train a feature search model and use the feature search model to complete the search processing for the target image to improve the calculation efficiency and accuracy of target search and tracking. The target search device can input the target sub-graph and the target image into the feature search model, and the feature search model can output the position information of the target feature information in the target image.

[0038] Please refer to Figure 3 , which is a schematic flowchart of the training of a feature search model provided by the embodiment of the present application. Before step S101, the following steps may be included: S201, construct an initial feature search model, obtain a sample sub-graph, a sample image, and the sample position information of the sample feature information of the sample sub-graph in the sample image.

[0039] Specifically, the target search and tracking device can construct an initial feature search model, which can include a backbone network, a search network, a target subgraph branch, and a feature database, and has the functions of performing feature extraction processing on the target subgraph and performing search processing on the target image. To further train the initial feature search model and optimize its feature extraction processing and search processing capabilities, the target search and tracking device can obtain sample subgraphs, sample images, and the sample position information of the sample feature information of the sample subgraph in the sample image. The sample subgraph is the sample data of the target subgraph for model training, the sample image is the sample data of the target image for model training, the sample feature information is the target feature information in the sample subgraph, and the sample position information is the target position information of the sample feature information in the sample image, which can be manually labeled by the user or relevant staff.

[0040] Optionally, the sample position information can be the position information of the sample annotation box in the sample image. The sample annotation box can contain the sample feature information, and the sample position information can include the upper left vertex coordinates of the sample annotation box, the length of the sample annotation box, and the width of the sample annotation box.

[0041] S202: Input the sample subgraph and the sample image into the initial feature search model to obtain the training position information output by the initial feature search model.

[0042] Specifically, input the sample subgraph and the sample image into the initial feature search model to obtain the training position information output by the initial feature search model. The initial feature search model can perform feature extraction processing on the target subgraph to obtain training feature information, and then obtain the training position information of the training feature information in the training image. The training position information is the prediction of the position of the training feature information in the training image by the initial feature search model, reflecting the current learning state, feature extraction processing ability, and search processing ability of the initial feature search model.

[0043] S203: Perform model training on the initial feature search model based on the training position information and the sample position information until the initial feature search model meets the training termination condition, and obtain the feature search model.

[0044] Specifically, the target search and tracking device can implement model training for the initial feature search model based on the training position information and the sample position information. By adjusting the parameters of the initial feature search model during the model training process, the loss value between the training position information and the sample position information is reduced, thereby improving the accuracy of the initial feature search model until the initial feature search model meets the training termination condition, and a feature search model is obtained. Among them, the training termination condition is used to determine whether the initial feature search model already has the feature extraction processing ability and the search processing ability, and whether it can accurately obtain the target position information. The training termination condition can be the initial setting of the target search and tracking device, or can be set by the user or relevant staff members.

[0045] Optionally, the loss value between the training position information and the sample position information can be the Mean Absolute Error (MAE), and the formula is as follows:

[0046] Where MAE is the loss value between the training position information and the sample position information, N is the number of sample images, is the sample position information, is the training position information.

[0047] Optionally, the initial feature search model has a multi-head output structure. Since there may be multiple target feature information during the user's use process, and there may be more than one target feature information in a target image, the initial feature search model can not only output the position information of the target feature information, but also output the category information corresponding to each target feature information. The category information can indicate whether each target feature information exists in the target image. One head can be used to obtain and output the position information, and one head can be used to obtain and output the category information. The multi-head output structure can enable the feature search model to perceive the features in the target image from multiple angles simultaneously. One head can focus on the position information, and the other head can focus on the category information, and it also enables the feature search model to reduce the dependence on a single feature and reduce the risk of overfitting.

[0048] Therefore, the sample data obtained by the target search and tracking device can also include the sample category information of the sample feature information. After inputting the sample sub-image and the sample image into the initial feature search model, the initial feature search model can not only output the training position information, but also output the training category information. The target search and tracking device can also perform model training on the initial feature search model based on the loss value between the training category information and the sample category information. Among them, the loss value between the training category information and the sample category information can be the Cross-Entropy Loss Function, and the formula is as follows:

[0049] Among them, is the loss value between the training category information and the sample category information, N is the number of sample images, M is the number of category information, is the sample category information, is the training category information.

[0050] Optionally, the training termination condition may include that the loss value drops below a preset loss value and remains for a first preset duration. The preset loss value is used to determine whether the prediction accuracy of the initial feature search model has reached the standard, and the first preset duration is used to determine whether the stability of the initial feature search model has reached the standard. Among them, the preset loss value and the first preset duration can be the initial settings of the target search and tracking device, or can be set by the user or relevant staff. For example, the preset loss value can be 0.05, and the first preset duration can be 100 rounds of model training. Among them, the loss value may include the loss value between the training position information and the sample position information and the loss value between the training category information and the sample category information.

[0051] Optionally, the training termination condition may include that the accuracy of the initial feature search model rises above a preset accuracy and remains for a second preset duration. The accuracy can be the proportion of the number of samples correctly predicted by the model to the total number of samples, and is an index for evaluating the overall performance of the model.

[0052] Optionally, the training termination condition may include that the precision and recall of the initial feature search model rise above a preset precision and recall and remain for a third preset duration. The precision and recall can consist of two parts: precision and recall. The precision can be the proportion of the sample pictures that actually have sample feature information among the sample pictures predicted by the model to have training feature information, which is used to reflect the accuracy of the model's prediction of positive classes. The recall is the proportion of the sample pictures that are correctly predicted by the model to have training feature information among the sample pictures that actually have sample feature information, which is used to reflect the coverage of the model for positive classes. The preset accuracy, the preset precision and recall, the second preset duration, and the third preset duration can be the initial settings of the target search and tracking device, or can be set by the user or relevant staff.

[0053] Optionally, the training termination condition may include that the number of model training rounds for the initial feature search model reaches a preset number of rounds. For example, the preset number of rounds can be 100 times. Using the number of training rounds as the training termination condition is very intuitive and easy to understand, does not require complex calculations or judgment conditions, and only needs to set the number of rounds at the beginning of training, making the management of the training process more concise.

[0054] Optionally, during the model training of the initial feature search model based on the training position information and the sample position information, the target search and tracking device can detect whether the initial feature search model meets the overfitting termination condition. If it meets the overfitting termination condition, the model training of the initial feature search model is paused and parameter adjustment prompt information for the initial feature search model is output. Since overfitting may occur during the model training process, pausing the model training in a timely manner can prevent the model from overlearning the noise and details in the training set, thereby maintaining good generalization ability. The parameter adjustment prompt information is used to prompt the user or relevant staff to check and modify the parameters of the initial feature search model, and continue the model training for the initial feature search model after the check and modification until the training termination condition is met to obtain the feature search model. Among them, the overfitting termination condition may include a sudden increase in the loss value and a large oscillation amplitude, as well as a sudden decrease in the precision-recall rate and accuracy and a large oscillation amplitude.

[0055] In the embodiment of the present application, an initial feature search model is constructed, a sample subgraph, a sample image, and the sample position information of the sample feature information of the sample subgraph in the sample image are obtained, the sample subgraph and the sample image are input into the initial feature search model, the training position information output by the initial feature search model is obtained, and the initial feature search model is trained based on the training position information and the sample position information until the initial feature search model meets the training termination condition to obtain the feature search model. The initial feature search model is trained specifically through the sample position information, and the model can stop training in a timely manner after reaching the termination condition, avoiding ineffective training iterations, saving training time, improving training efficiency, and combining the overfitting termination condition to prevent the model from overlearning the noise and details in the training set, thereby maintaining good generalization ability.

[0056] Please refer to Figure 4 , which provides a schematic flowchart of a feature extraction process. Step S101 may include the following steps: S301, obtain the target subgraph input by the user.

[0057] Specifically, the user can input the target subgraph containing the target to be tracked into the target search and tracking device. The target search and tracking device can obtain the target subgraph input by the user. The target subgraph can be obtained by the user cutting the target image according to the tracking target.

[0058] Optionally, the user can input an image and the target annotation boxes marked by the user. The target annotation boxes contain the targets that the user needs to track. The target search and tracking device crops the input image of the user based on the target annotation boxes to obtain target sub-images. The input image of the user may contain one or more target annotation boxes, and each target annotation box corresponds to a target sub-image.

[0059] S302: Input the target sub-image into the target sub-image branch of the feature search model, and perform feature extraction processing on the target sub-image to obtain target feature information.

[0060] Specifically, the target search and tracking device can input the target sub-image into the sub-image branch of the feature search model. The sub-image branch can be used to perform extraction processing on the target sub-image to obtain target feature information. Among them, the target sub-image can be one or more pictures, and each target sub-image can contain a target feature information. By specifically performing feature extraction processing on the target sub-image through the target sub-image branch, the target sub-image branch can be more focused on the feature information of the target sub-image and avoid being interfered by irrelevant information. This targeted feature extraction method helps to extract more representative and distinguishable target feature information and improve the quality and efficiency of feature extraction.

[0061] S303: Store the target feature information in the feature database of the feature search model.

[0062] Specifically, after the target sub-image obtains the target feature information, it can send the target feature information to the feature database of the feature search model for storage, which is convenient for directly obtaining the target feature information from the feature database during subsequent search processing. The user can input different target sub-images according to their own needs, and the target search and tracking device can extract and store the corresponding target feature information. This personalized customization method can better meet the target search and tracking needs of users in specific scenarios, such as searching and tracking specific personnel or vehicles in security monitoring, and searching and tracking specific defects in industrial inspection.

[0063] Optionally, there may be multiple target feature information, so the target search device can determine the category information corresponding to the target feature information and associate and store the category information with the target feature information in the feature database.

[0064] In an embodiment of the present application, a target sub-graph input by a user is obtained, and the target sub-graph is input into the target sub-graph branch of a feature search model. Feature extraction processing is performed on the target sub-graph to obtain target feature information, and the target feature information is stored in the feature database of the feature search model. By specifically performing feature extraction processing on the target sub-graph through the target sub-graph branch, the quality and efficiency of target feature information extraction are improved. Moreover, the target feature information input by the user is extracted and stored, which can meet the personalized needs of the user and improve the adaptability of the model in different scenarios.

[0065] Please refer to Figure 5 , which is a schematic flowchart of a search processing provided by an embodiment of the present application. Step S102 may include the following steps: S401, Obtain a target image input by a user, and obtain target feature information in the feature database.

[0066] Specifically, a user may input a target image into a target search and tracking device. The target image may be one or more images. For example, the user may input multiple target images, and the target search and tracking device may perform search processing on all the target images. The user may also input video data, and the target image is each frame of the video data, thereby realizing the tracking of the target in the video. The target search and tracking device may obtain the target image input by the user and may obtain the target feature information in the feature database.

[0067] Optionally, the target search and tracking device may obtain all the target feature information in the feature database and perform search processing on all the target feature information on the target image.

[0068] S402, Input the target image into the search network of the feature search model, perform correlation convolution processing on the target image and the target feature information, and determine whether there is target feature information in the target image.

[0069] Specifically, the target search and tracking device may input the target image into the search network of the feature search model. The search network may perform correlation convolution processing (Correlation Convolution) on the target image and the target feature information, thereby searching for the target feature information in the target image and determining whether there is target feature information in the target image. The correlation convolution processing includes a convolution operation and a correlation calculation. The convolution operation may slide a convolution kernel on the target image, perform dot multiplication and summation to extract features. The correlation calculation is a method for measuring the similarity between two data. By calculating the correlation feature tensor between the image feature information in the target image and the target feature information, and determining whether there is target feature information according to the correlation feature tensor.

[0070] Optionally, the target search and tracking device inputs the target image into the search network of the feature search model, and then controls the search network to perform a convolution operation on the target image using the target feature information as a convolution kernel, obtaining associated feature information that matches the target feature information. The search network can use the target feature information as a convolution kernel and then slide the convolution kernel on the target image, thereby calculating the associated feature tensor between the target image and the image feature information of the target image. Based on the feature search model with a multi-head output structure, classification and regression methods are respectively used to obtain the category information indicating whether the target feature information exists and the position information of the target feature information. For example, the search network can use the target feature information as a convolution kernel and then slide the convolution kernel on the target image, thereby calculating the similarity between the target image and the image feature information of the target image. The image feature information with a similarity greater than the preset similarity to the target feature information is used as the image feature information that matches the target feature information. If there is associated feature information that matches the target feature information in the target image, it indicates that the target feature information exists in the target image, and the position information of the associated feature information in the target image is the position information of the target feature information in the target image.

[0071] In the embodiment of the present application, the target image input by the user is obtained, the target feature information is obtained from the feature database, the target image is input into the search network of the feature search model, and an associated convolution process is performed on the target image and the target feature information to determine whether the target feature information exists in the target image. Through the associated convolution process, precise matching and rapid positioning of the target feature information can be achieved, reducing missed detection cases and improving the search efficiency.

[0072] Please refer to Figure 6 , which provides a schematic flowchart for obtaining position information in the embodiment of the present application. Step S103 may include the following steps: S501, if there is associated feature information that matches the target feature information, it is determined that the target feature information exists in the target image, and the category information of the target feature information is obtained.

[0073] Specifically, if it is detected that there is associated feature information that matches the target feature information in the target image, it can be determined that the target feature information exists in the target image, and the category information of the target feature information is obtained. The category information can be used to indicate whether the target feature information exists in the target image. Therefore, if the target feature information exists, the category information can be 1, and conversely, if the target feature information does not exist, the category information can be 0.

[0074] Optionally, since multiple target feature information can be stored in the feature database, in order to distinguish which target feature information each category of information belongs to, the target search and tracking device can sort the target feature information in the feature database according to a preset order. The preset order can be the initial setting of the target search and tracking device or can be set by relevant staff. For example, it can be sorted according to storage time or the first letter, or sorted by relevant staff. Then, the category information of each target feature information can be output according to the preset order. For example, if there are three target feature information in the feature database, and the preset order of these three target feature information is target feature information A, target feature information B, and target feature information C, and then the target search and tracking device detects that target feature information A and target feature information C exist in the target image, and target feature information B does not exist, then the category information "1, 1, 0" can be output according to the preset order. Optionally, in order to further distinguish the target feature information, relevant staff or the target search and tracking device can set a category flag for each target feature information. The category flag can be a number or a name, which can be used to distinguish each target feature information and can also be used to briefly describe each target feature information. When saving the target feature information to the feature database, the category flag can be associated and stored with the target feature information in the feature database. If there is associated feature information that matches the target feature information, it is determined that the target feature information exists in the target image, and the category flag of the target feature information is obtained, and category information is generated based on the category flag. That is, the category information can include the category flag, which is used to indicate whether the target feature information corresponding to the category flag exists in the target image. For example, if there are three target feature information in the feature database, and the category identifier of target feature information A is "A", the category identifier of target feature information B is "B", and the category identifier of target feature information C is "C", and then the target search and tracking device detects that target feature information A and target feature information C exist in the target image, and target feature information B does not exist, then the category information "A, 1", "B, 0", and "C, 1" can be obtained.

[0075] S502. Obtain the position information of the associated feature information in the target image.

[0076] Specifically, the position information of the associated feature information in the target image is the position information of the target feature information in the target image. Therefore, the target search and tracking device can control the search network to obtain the position information of the associated feature information in the target image.

[0077] Optionally, the position information can be the coordinate information of the associated feature information or the position information of the annotation box containing the associated feature information in the target image. For example, the position information can include the upper left vertex coordinates of the annotation box, the length of the annotation box, and the width of the annotation box.

[0078] Optionally, if there is no associated feature information that matches the target feature information, it is determined that the target feature information does not exist in the target image, and the category information of the target feature information is obtained. At this time, the category information can indicate that the target feature information does not exist in the target image. For example, the category information can be 0. Since the feature search model has a multi-head output structure, if it is detected that the target feature information does not exist in the target image, there is no need to calculate the position information anymore.

[0079] S503, output the category annotation and the position information.

[0080] Specifically, the target search and tracking device can control the search network to output the category annotation and the position information, and the target search and tracking device can generate a bounding box in the target image according to the position information and display the target image with the bounding box to the user, so as to intuitively display the target in the target image to the user, realize the search and tracking of the target, and by obtaining the category information of the target feature information, it helps to improve the accuracy of the target search and tracking, avoid confusing targets of different categories, improve the reliability of the search results, reduce the situation of missed detection and false detection, and also enable the target search and tracking device to realize the search and tracking of multiple targets, and be more adaptable to diverse application scenarios.

[0081] In the embodiment of the present application, if there is associated feature information that matches the target feature information, it is determined that the target feature information exists in the target image, the category information of the target feature information is obtained, the position information of the associated feature information in the target image is obtained, and the category annotation and the position information are output. By obtaining the category information of the target feature information, it helps to improve the accuracy of the target search and tracking, avoid confusing targets of different categories, improve the reliability of the search results, and also enable the target search and tracking device to realize the search and tracking of multiple targets, and be more adaptable to diverse application scenarios.

[0082] The following will be combined with the attached Figure 7 - attached Figure 8 , and the target search and tracking device provided in the embodiment of the present application will be introduced in detail. It should be noted that the target search and tracking device in the attached Figure 7 - attached Figure 8 is used to execute the method of the embodiment shown in the present application. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown. For the specific technical details not disclosed, please refer to the embodiment shown in the present application Figures 1 - 6 shown. Figures 1 - 6 shown in the embodiment.

[0083] Please refer to Figure 7, which shows a schematic structural diagram of a target search and tracking device provided by an exemplary embodiment of the present application. The target search and tracking device can be implemented as all or part of the device through software, hardware, or a combination of both. The device 1 includes a target feature acquisition unit 11, a target feature search unit 12, and a position information determination unit 13.

[0084] The target feature acquisition unit 11 is configured to acquire a target sub-graph input by a user, and perform feature extraction processing on the target sub-graph to obtain target feature information; The target feature search unit 12 is configured to acquire the target image input by the user, and perform a search process on the target feature information in the target image; The position information determination unit 13 is configured to, if the target feature information exists in the target image, acquire the position information of the target feature information in the target image.

[0085] In this embodiment, a target sub-graph input by a user is acquired, feature extraction processing is performed on the target sub-graph to obtain target feature information, the target image input by the user is acquired, a search process is performed on the target feature information in the target image, and if the target feature information exists in the target image, the position information of the target feature information in the target image is acquired. By searching for the target feature in the target sub-graph in the target image to determine the position of the target feature, and extracting feature information to achieve the purpose of tracking the target feature with less computing power, it is applicable to more applications and reduces the computing power requirements for the electronic device.

[0086] Please refer to Figure 8 , which shows a schematic structural diagram of a target search and tracking device provided by an exemplary embodiment of the present application. The target search and tracking device can be implemented as all or part of the device through software, hardware, or a combination of both. The device 1 includes a model training unit 14, a target feature acquisition unit 11, a target feature search unit 12, and a position information determination unit 13.

[0087] The model training unit 14 is configured to construct an initial feature search model, acquire a sample sub-graph, a sample image, and sample position information of the sample feature information of the sample sub-graph in the sample image; Input the sample sub-graph and the sample image into the initial feature search model to obtain training position information output by the initial feature search model; Based on the training position information and the sample position information, perform model training on the initial feature search model until the initial feature search model meets the training termination condition, and obtain a feature search model.

[0088] Optionally, the model training unit 14 is specifically configured to perform model training on the initial feature search model based on the training position information and the sample position information until the initial feature search model meets the overfitting termination condition, and output parameter adjustment prompt information for the initial feature search model.

[0089] The target feature acquisition unit 11 is configured to acquire a target subgraph input by a user, and perform feature extraction processing on the target subgraph to obtain target feature information; Optionally, the target feature acquisition unit 11 is specifically configured to acquire the target subgraph input by the user; Input the target subgraph into the target subgraph branch of the feature search model, and perform feature extraction processing on the target subgraph to obtain target feature information; Store the target feature information in the feature database of the feature search model.

[0090] The target feature search unit 12 is configured to acquire a target image input by the user, and perform a search process on the target feature information in the target image; Optionally, the target feature search unit 12 is specifically configured to acquire the target image input by the user, and acquire the target feature information from the feature database; Input the target image into the search network of the feature search model, perform associated convolution processing on the target image and the target feature information, and determine whether the target feature information exists in the target image.

[0091] Optionally, the target feature search unit 12 is specifically configured to input the target image into the search network of the feature search model; Control the search network to use the target feature information as a convolution kernel to perform a convolution operation on the target image, and obtain associated feature information that matches the target feature information.

[0092] The position information determination unit 13 is configured to, if the target feature information exists in the target image, acquire the position information of the target feature information in the target image.

[0093] Optionally, the position information determination unit 13 is specifically configured to, if there is associated feature information that matches the target feature information, determine that the target feature information exists in the target image, and acquire the category information of the target feature information; Acquire the position information of the associated feature information in the target image; Output the category information and the position information.

[0094] In this embodiment, an initial feature search model is constructed, and sample subgraphs, sample images, and sample position information of the sample feature information of the sample subgraphs in the sample images are obtained. The sample subgraphs and sample images are input into the initial feature search model to obtain the training position information output by the initial feature search model. The initial feature search model is trained based on the training position information and the sample position information until the initial feature search model meets the training termination condition, and a feature search model is obtained. By training the initial feature search model specifically with the sample position information and enabling the model to stop training in a timely manner after reaching the termination condition, invalid training iterations are avoided, training time is saved, training efficiency is improved, and overfitting termination conditions can be combined to prevent the model from overlearning the noise and details in the training set, thus maintaining good generalization ability. The target subgraph input by the user is obtained, and the target subgraph is input into the target subgraph branch of the feature search model. Feature extraction processing is performed on the target subgraph to obtain target feature information, and the target feature information is stored in the feature database of the feature search model. By specifically performing feature extraction processing on the target subgraph through the target subgraph branch, the quality and efficiency of target feature information extraction are improved, and the target feature information input by the user is extracted and stored, which can meet the personalized needs of the user and improve the adaptability of the model in different scenarios. The target image input by the user is obtained, the target feature information is obtained from the feature database, and the target image is input into the search network of the feature search model. Correlation convolution processing is performed on the target image and the target feature information to determine whether the target feature information exists in the target image. Precise matching and fast positioning of the target feature information can be achieved through correlation convolution processing, reducing missed detection cases and improving search efficiency. If there is associated feature information that matches the target feature information, it is determined that the target feature information exists in the target image, the category information of the target feature information is obtained, the position information of the associated feature information in the target image is obtained, and the category annotation and position information are output. By obtaining the category information of the target feature information, it helps to improve the accuracy of target search and tracking, avoid confusing targets of different categories, improve the reliability of search results, and enable the target search and tracking device to achieve search and tracking of multiple targets, making it more adaptable to diverse application scenarios.

[0095] It should be noted that when the target search and tracking device provided in the above embodiment executes the target search and tracking method, only the above division of each functional module is used as an example. In practical applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the target search and tracking device provided in the above embodiment and the target search and tracking method embodiment belong to the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.

[0096] The serial numbers of the embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0097] The embodiments of the present application also provide a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the target search and tracking method of the embodiments as described above Figures 1 - 6 The specific execution process can be referred to Figures 1 - 6 the specific description of the embodiments shown, and will not be elaborated here.

[0098] The present application also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor to perform the target search and tracking method of the embodiments as described above Figures 1 - 6 The specific execution process can be referred to Figures 1 - 6 the specific description of the embodiments shown, and will not be elaborated here.

[0099] Please refer to Figure 9 , which shows a block diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. The electronic device in the present application may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected through the bus 150.

[0100] The processor 110 may include one or more processing cores. The processor 110 uses various interfaces and lines to connect various parts within the entire electronic device, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120, it executes various functions of the terminal 100 and processes data. Optionally, the processor 110 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 110 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user pages, and application programs, etc.; the GPU is responsible for rendering and drawing the display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 110 and may be implemented separately through a communication chip.

[0101] The memory 120 may include a Random Access Memory (RAM), and may also include a Read-Only Memory (ROM). Optionally, the memory 120 includes a Non-Transitory Computer-Readable Storage Medium. The memory 120 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above method embodiments, etc. The operating system may be an Android system, including a system developed based on the Android system in depth, an IOS system developed by Apple Inc., including a system developed based on the IOS system in depth, or other systems.

[0102] The memory 120 can be divided into an operating system space and a user space. The operating system runs in the operating system space, and native and third-party application programs run in the user space. In order to ensure that different third-party application programs can achieve better running effects, the operating system allocates corresponding system resources for different third-party application programs. However, there are also differences in the system resource requirements of different application scenarios in the same third-party application program. For example, in the local resource loading scenario, the third-party application program has a higher requirement for the disk read speed; in the animation rendering scenario, the third-party application program has a higher requirement for the GPU performance. The operating system and the third-party application program are independent of each other, and the operating system often cannot timely perceive the current application scenario of the third-party application program, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application program.

[0103] In order to enable the operating system to distinguish the specific application scenarios of third-party application programs, it is necessary to establish data communication between the third-party application programs and the operating system, so that the operating system can obtain the current scenario information of the third-party application programs at any time, and then perform targeted system resource adaptation based on the current scenario.

[0104] Among them, the input device 130 is used to receive input instructions or data. The input device 130 includes, but is not limited to, a keyboard, a mouse, a camera, a microphone or a touch device. The output device 140 is used to output instructions or data. The output device 140 includes, but is not limited to, a display device and a speaker, etc. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are a touch display screen.

[0105] The touch display screen can be designed as a full-screen, curved screen or irregular-shaped screen. The touch display screen can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregular-shaped screen and a curved screen. The embodiments of the present application do not limit this.

[0106] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above drawings does not limit the electronic device. The electronic device may include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements. For example, the electronic device also includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, a Wireless Fidelity (WiFi) module, a power supply, a Bluetooth module, etc., which will not be elaborated here.

[0107] In Figure 9 In the electronic device shown, the processor 110 can be used to call the target search and tracking application program stored in the memory 120 and specifically perform the following operations: Obtain the target sub-graph input by the user, and perform feature extraction processing on the target sub-graph to obtain target feature information; Obtain the target image input by the user, and perform search processing on the target feature information in the target image; If the target feature information exists in the target image, obtain the position information of the target feature information in the target image.

[0108] In one embodiment, before the processor 110 executes obtaining the target sub-graph input by the user and performing feature extraction processing on the target sub-graph to obtain target feature information, the following operations are also performed: Construct an initial feature search model, obtain a sample sub-graph, a sample image, and the sample position information of the sample feature information of the sample sub-graph in the sample image; Input the sample sub-graph and the sample image into the initial feature search model to obtain the training position information output by the initial feature search model; Based on the training position information and the sample position information, perform model training on the initial feature search model until the initial feature search model meets the training termination condition to obtain a feature search model.

[0109] In one embodiment, when the processor 110 executes obtaining the target sub-graph input by the user and performing feature extraction processing on the target sub-graph to obtain target feature information, the following operations are specifically performed: Obtain the target sub-graph input by the user; Input the target sub - graph into the target sub - graph branch of the feature search model, and perform feature extraction processing on the target sub - graph to obtain target feature information; Store the target feature information in the feature database of the feature search model.

[0110] In one embodiment, when the processor 110 executes to obtain the target image input by the user and perform a search process on the target feature information in the target image, the following operations are specifically performed: Obtain the target image input by the user, and obtain the target feature information in the feature database; Input the target image into the search network of the feature search model, perform associated convolution processing on the target image and the target feature information, and determine whether the target feature information exists in the target image.

[0111] In one embodiment, when the processor 110 executes to input the target image into the search network of the feature search model and perform associated convolution processing on the target image and the target feature information, the following operations are specifically performed: Input the target image into the search network of the feature search model; Control the search network to use the target feature information as a convolution kernel to perform a convolution operation on the target image, and obtain associated feature information that matches the target feature information.

[0112] In one embodiment, when the processor 110 executes if the target feature information exists in the target image, then obtain the position information of the target feature information in the target image, the following operations are specifically performed: If there is associated feature information that matches the target feature information, determine that the target feature information exists in the target image, and obtain the category information of the target feature information; Obtain the position information of the associated feature information in the target image; Output the category information and the position information.

[0113] In one embodiment, when the processor 110 executes to perform model training on the initial feature search model based on the training position information and the sample position information until the initial feature search model meets the training termination condition and obtain the feature search model, the following operations are specifically performed: Perform model training on the initial feature search model based on the training position information and the sample position information until the initial feature search model meets the over - fitting termination condition, and output parameter adjustment prompt information for the initial feature search model.

[0114] In this embodiment, an initial feature search model is constructed. Sample subgraphs, sample images, and sample position information of the sample feature information of the sample subgraphs in the sample images are obtained. The sample subgraphs and the sample images are input into the initial feature search model to obtain the training position information output by the initial feature search model. The initial feature search model is trained based on the training position information and the sample position information until the initial feature search model meets the training termination condition, and a feature search model is obtained. By using the sample position information to specifically train the initial feature search model, the model can stop training in time after reaching the termination condition, avoiding ineffective training iterations, saving training time, and improving training efficiency. In addition, overfitting termination conditions can be combined to prevent the model from overlearning the noise and details in the training set, thereby maintaining good generalization ability. The target subgraph input by the user is obtained, and the target subgraph is input into the target subgraph branch of the feature search model. Feature extraction processing is performed on the target subgraph to obtain target feature information, and the target feature information is stored in the feature database of the feature search model. By specifically performing feature extraction processing on the target subgraph through the target subgraph branch, the quality and efficiency of target feature information extraction are improved. Moreover, the target feature information input by the user is extracted and stored, which can meet the personalized needs of the user and improve the adaptability of the model in different scenarios. The target image input by the user is obtained, the target feature information is obtained from the feature database, and the target image is input into the search network of the feature search model. Correlation convolution processing is performed on the target image and the target feature information to determine whether target feature information exists in the target image. Through correlation convolution processing, precise matching and rapid positioning of the target feature information can be achieved, reducing missed detection cases and improving search efficiency. If there is associated feature information that matches the target feature information, it is determined that target feature information exists in the target image, the category information of the target feature information is obtained, the position information of the associated feature information in the target image is obtained, and the category annotation and the position information are output. By obtaining the category information of the target feature information, it helps to improve the accuracy of target search and tracking, avoid confusing targets of different categories, improve the reliability of search results, and enable the target search and tracking device to implement search and tracking for multiple targets, making it more suitable for diverse application scenarios.

[0115] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, or the like.

[0116] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

[0117] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the target subgraphs, target feature information, target images, etc. involved in this specification are all obtained under full authorization.

Claims

1. A target search and tracking method, characterized in that: The method comprises: Obtaining a target subgraph input by a user, and performing feature extraction processing on the target subgraph to obtain target feature information; Acquire the target image input by the user, and perform search processing on the target feature information in the target image; If the target feature information exists in the target image, the position information of the target feature information in the target image is obtained.

2. The method according to claim 1, characterized in that Before obtaining the target subgraph input by the user and performing feature extraction processing on the target subgraph to obtain target feature information, the method further includes: Constructing an initial feature search model, obtaining a sample sub-graph, a sample image, and sample position information of sample feature information of the sample sub-graph in the sample image; Inputting the sample sub-image and the sample image into the initial feature search model to obtain training position information output by the initial feature search model; The initial feature search model is trained based on the training position information and the sample position information until the initial feature search model meets the training termination condition, thereby obtaining a feature search model.

3. The method according to claim 2, characterized in that The step of obtaining a target subgraph input by a user and performing feature extraction processing on the target subgraph to obtain target feature information includes: Obtaining the target subgraph input by the user; Inputting the target subgraph into the target subgraph branch of the feature search model, performing feature extraction processing on the target subgraph to obtain target feature information; The target feature information is stored in a feature database of the feature search model.

4. The method according to claim 2, characterized in that: The step of acquiring the target image input by the user and searching for the target feature information in the target image includes: Acquire the target image input by the user, and acquire the target feature information in a feature database; The target image is input into the search network of the feature search model, and an associated convolution process is performed on the target image and the target feature information to determine whether the target feature information exists in the target image.

5. The method according to claim 4, characterized in that The step of inputting the target image into the search network of the feature search model and performing associated convolution processing on the target image and the target feature information includes: Inputting the target image into the search network of the feature search model; The search network is controlled to perform a convolution operation on the target image using the target feature information as a convolution kernel to obtain associated feature information consistent with the target feature information.

6. The method according to claim 5, characterized in that If the target feature information exists in the target image, obtaining the position information of the target feature information in the target image includes: If there is associated feature information that matches the target feature information, determining that the target feature information exists in the target image, and acquiring category information of the target feature information; Acquire location information of the associated feature information in the target image; The category information and the position information are output.

7. The method according to claim 2, characterized in that The performing model training on the initial feature search model based on the training position information and the sample position information until the initial feature search model meets the training termination condition to obtain the feature search model includes: The initial feature search model is trained based on the training position information and the sample position information until the initial feature search model meets an overfitting termination condition, and parameter adjustment prompt information for the initial feature search model is output.

8. A target search and tracking device, characterized in that: The device comprises: A target feature acquisition unit, used to acquire a target sub-graph input by a user, and perform feature extraction processing on the target sub-graph to obtain target feature information; A target feature search unit, used to obtain the target image input by the user, and perform a search process on the target feature information in the target image; A position information determining unit is used to obtain position information of the target feature information in the target image if the target feature information exists in the target image.

9. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method steps as claimed in any one of claims 1 to 7.

10. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps as claimed in any one of claims 1 to 7.