Target detection method and device, terminal equipment and storage medium
By combining image and text information for in-depth analysis and processing, the problem of only limited category information in the existing technology is solved, the target detection of new categories is realized, the application scope of object detection technology is expanded, and the diversified detection needs of SaaS platforms are met.
Patent Information
- Application Number
- CN202510238420.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-06
AI Technical Summary
Existing target detection technology can only identify limited category information and cannot effectively detect the target content of new categories in a timely manner, resulting in a small application scope and cannot meet the target detection needs of diversified data content in the SaaS platform field.
By combining image information and text information for deep analysis, the preset image feature extraction model, text feature extraction model, feature enhancement vector and fusion feature processing model are used to generate target image feature information and target text feature information, and finally identify specific targets in the image through the feature decoding model and feature analysis model.
It can identify specific targets in the image without creating a training set, shorten the detection cycle, expand the application scope of object detection technology, and meet the timeliness and reliability requirements of SaaS platform for object detection of new category data.
Smart Images

Figure CN120107556A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology, and in particular, relates to a target detection method, apparatus, terminal device and storage medium. Background Art
[0002] Object detection is an important key component technology in computer vision, which is generally used to locate and identify specific targets in a given image. In the field of SaaS platforms, object detection technology is usually used to identify specific target content to meet customers' target point detection needs.
[0003] In the existing technology, target detection technology only identifies specific target content based on image information, which usually requires a lot of manpower and material resources to collect and label image information data, and to create training sets for training existing deep learning models, and then use the trained deep learning models to realize the detection needs of specific targets in the image.
[0004] However, in the SaaS platform field, customers' target detection content is ever-changing, which makes it difficult to collect image information and create data sets. As a result, existing technologies cannot meet the detection needs of many target categories in the SaaS platform in a timely manner, which will lead to customer loss. Summary of the invention
[0005] In view of this, the embodiments of the present application provide a target detection method, apparatus, terminal device and storage medium, aiming to solve the problem that the target detection technology can only realize the identification and positioning of limited category information, and cannot effectively detect the target content of new categories in time, which makes the application scope of target detection small and difficult to expand, and cannot meet the target detection needs of diverse categories of data content in the SaaS platform field.
[0006] A first aspect of an embodiment of the present application provides a target detection method, including:
[0007] Obtaining image information and text information of the target to be detected;
[0008] Obtaining initial image feature information and initial text feature information according to the target image information to be detected, the target text information to be detected, a preset image feature extraction model, and a preset text feature extraction model;
[0009] Obtaining target image feature information and target text feature information according to the initial image feature information, the initial text feature information, a preset feature enhancement vector, and a preset fusion feature processing model;
[0010] The target image information is obtained according to the target image feature information, the target text feature information, a preset feature decoding model and a preset feature parsing model.
[0011] A second aspect of an embodiment of the present application provides a target detection device, including:
[0012] The target information acquisition module is used to acquire the image information and text information of the target to be detected;
[0013] An initial feature information generation module, used to obtain initial image feature information and initial text feature information according to the target image information to be detected, the target text information to be detected, a preset image feature extraction model and a preset text feature extraction model;
[0014] a target feature information generating module, configured to obtain target image feature information and target text feature information according to the initial image feature information, the initial text feature information, a preset feature enhancement vector and a preset fusion feature processing model; and
[0015] The target image generation module is used to obtain the target image information according to the target image feature information, the target text feature information, a preset feature decoding model and a preset feature parsing model.
[0016] A third aspect of an embodiment of the present application provides a terminal device, which includes a memory and a processor, wherein the memory stores a computer program that can be executed on the processor, and when the processor executes the computer program, the steps of the target detection method described in the first aspect above are implemented.
[0017] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, comprising: a computer program stored therein, wherein when the computer program is executed by a processor, the steps of the target detection method described in the first aspect above are implemented.
[0018] Compared with the prior art, the embodiments of the present application have the following beneficial effects: by combining image information and text information for deep analysis processing, specific targets in the image can be identified based on the text description content without the need to create a training set for training the computing model, thereby avoiding the detection scope being limited to predefined positioning or identification target categories, thereby greatly shortening the detection cycle, and being able to handle more diverse and fine-grained target detection tasks, effectively expanding the application scope of target detection technology, and meeting the SaaS platform's requirements for timeliness and reliability in target detection of new category data. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0020] Figure 1 This is a schematic diagram of the implementation process of the target detection method provided in Example 1 of the present application;
[0021] Figure 2 This is a schematic diagram of the implementation process of the target detection method provided in Example 2 of the present application;
[0022] Figure 3 This is a schematic diagram of the implementation process of the target detection method provided in Example 3 of the present application;
[0023] Figure 4 This is a schematic diagram of the implementation process of the target detection method provided in Example 4 of the present application;
[0024] Figure 5 This is a schematic diagram of the implementation process of the target detection method provided in Example 5 of the present application;
[0025] Figure 6 This is a schematic diagram of the implementation flow of the target detection method provided in Example 6 of the present application;
[0026] Figure 7 is a schematic diagram of the structure of a target detection device provided in an embodiment of the present application;
[0027] Figure 8 It is a schematic diagram of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0029] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.
[0030] Figure 1 The following is a flowchart of the target detection method according to the first embodiment of the present invention.
[0031] Step S101, obtaining image information and text information of a target to be detected.
[0032] In this embodiment, the target to be detected can be a specific object, such as shoes, clothes, eyes, etc., or an animal or a person. The image information of the target to be detected can be taken by a surveillance camera. The text information of the target to be detected can be manually input. It can be understood that the target to be detected is part of the information in the image information of the target to be detected, and the target to be detected in the image information of the target to be detected needs to be identified by the target detection method.
[0033] Step S102, obtaining initial image feature information and initial text feature information according to the to-be-detected target image information, the to-be-detected target text information, a preset image feature extraction model, and a preset text feature extraction model.
[0034] In this embodiment, the preset image feature extraction model may be Swin Transformer, which is used to extract features of image information; the preset text feature extraction model may be BERT, which is used to extract semantic features of text information. The target image information to be detected may be used as input information of the image feature extraction model, and then the image feature extraction model is used to extract features of the target image information to be detected, thereby using the semantic features of the target image to be detected, that is, the initial image feature information, as output information of the image feature extraction model. The target text information to be detected may be used as input information of the text feature extraction model, and then the text feature extraction model is used to extract features of the target text information to be detected, thereby using the semantic features of the target text to be detected, that is, the initial text feature information, as output information of the text feature extraction model.
[0035] Step S103, obtaining target image feature information and target text feature information according to the initial image feature information, the initial text feature information, a preset feature enhancement vector and a preset fusion feature processing model.
[0036] In this embodiment, the preset feature enhancement vector may be a query vector, a key vector, and a value vector in the attention mechanism, which is used to perform enhancement processing and fusion calculation on the initial image feature information and the initial text feature information. The preset fusion feature processing model may be a feedforward neural network, which is used to perform further in-depth processing on the fused features, thereby outputting in-depth processed image features and text features, i.e., target image feature information and target text feature information, to facilitate subsequent decoding processing to resolve the position of the target to be detected in the image information.
[0037] Step S104, obtaining target image information according to the target image feature information, the target text feature information, a preset feature decoding model and a preset feature parsing model.
[0038] In this embodiment, the preset feature decoding model can be set based on the self-attention mechanism. Specifically, it can be composed of one or more attention modules for processing image features, one or more attention modules for processing the fusion features of image features and text features, and one or more attention modules for processing text features, and is used to decode the target image feature information and the target text feature information. The preset feature parsing model can be a feedforward neural network, which is used to parse the feature information to obtain the position information of the target to be detected in the image information. The target image feature information and the target text feature information can be processed by the feature decoding model first, and then processed by the feature parsing model, so as to output the position information of the target to be detected in the image information, that is, the target image information, so as to achieve target detection.
[0039] In this embodiment, optionally, the fused and enhanced text features and image features are input into the decoder. Since the fused and enhanced image features have text guidance information added, in order to further strengthen the fused features, the image features are self-attention operated, and then the text features are combined to perform cross-self-attention operations on the text. The query capability of the computing model is improved through the cross-attention layer of image and text, and finally the target location information is output through a feedforward neural network.
[0040] The target detection method provided in the embodiment of the present application combines image information and text information for deep analysis and processing. It does not need to prepare a training set for training the computing model. It can identify specific targets in the image based on the text description content, thereby avoiding the detection scope being limited to predefined positioning or identification target categories, thereby greatly shortening the detection cycle and being able to handle more diverse and fine-grained target detection tasks, effectively expanding the application scope of target detection technology, and meeting the SaaS platform's requirements for timeliness and reliability in target detection of new category data.
[0041] Figure 2 The following is a flowchart of the target detection method provided in the second embodiment of the present application, which differs from the first embodiment in that:
[0042] The preset feature enhancement vector includes a preset image query vector, a preset image key vector, a preset image value vector, a preset text query vector, a preset text key vector and a preset text value vector;
[0043] The step S103 specifically includes:
[0044] Step S201, calculating image feature intermediate variable information according to the initial image feature information, a preset image query vector, a preset image key vector and a preset image value vector.
[0045] In this embodiment, the preset image query vector, the preset image key vector, the preset image value vector, the preset text query vector, the preset text key vector and the preset text value vector may be manually set or obtained after training certain specific computing models. The initial image feature information may be firstly subjected to dot multiplication calculation with the preset image query vector, the preset image key vector and the preset image value vector, and the calculation result may obtain the image query feature information, the image key feature information and the image value feature information, and the image query feature information, the image key feature information and the image value feature information may be collectively referred to as image feature intermediate variable information.
[0046] In this embodiment, optionally, image feature intermediate variable information may be calculated based on a deformable attention mechanism according to initial image feature information, a preset image query vector, a preset image key vector, and a preset image value vector.
[0047] Step S202, calculating text feature intermediate variable information according to the initial text feature information, a preset text query vector, a preset text key vector, and a preset text value vector.
[0048] In this embodiment, the initial text feature information may be firstly subjected to dot multiplication calculations with a preset text query vector, a preset text key vector, and a preset text value vector, respectively. The calculation results may be text query feature information, text key feature information, and text value feature information. The text query feature information, text key feature information, and text value feature information may be collectively referred to as text feature intermediate variable information.
[0049] In this embodiment, optionally, the text feature intermediate variable information may be calculated based on the self-attention mechanism according to the initial text feature information, a preset text query vector, a preset text key vector, and a preset text value vector.
[0050] Step S203, calculating the image-text cross-attention feature information and the text-image cross-attention feature information according to the image feature intermediate variable information and the text feature intermediate variable information.
[0051] In this embodiment, the intermediate variables in the image feature intermediate variable information and the text feature intermediate variable information can be cross-multiplied respectively, and the multiplication results are used as the image-text cross-attention feature information and the text-image cross-attention feature information. Among them, the image-text cross-attention feature information represents the cross-attention weight information mapped from the image to the text, which can be obtained by multiplying a small number of variables in the text feature intermediate variable information with a large number of variables in the image feature intermediate variable information; the text-image cross-attention feature information represents the cross-attention weight information mapped from the text to the image, which can be obtained by multiplying a large number of variables in the text feature intermediate variable information with a small number of variables in the image feature intermediate variable information.
[0052] Step S204, obtaining target image feature information and target text feature information according to the image-text cross-attention feature information, the text-image cross-attention feature information and a preset fusion feature processing model.
[0053] In this embodiment, the preset fusion feature processing model can be a feedforward neural network. The image-text cross-attention feature information and the text-image cross-attention feature information can be used as the input of the fusion feature processing model, so that through the calculation and processing of the fusion feature processing model, the target image feature information and the target text feature information are output, which are used to represent the updated image feature information and the updated text feature information, that is, the target image feature information and the target text feature information. It can be understood that the image features and the text features have been fused and calculated in the previous steps of calculating the image-text cross-attention feature information and the text-image cross-attention feature information, so the updated image feature information, that is, the target image feature information, already carries part of the text feature information, and the updated text feature information, that is, the target text feature information, already carries part of the image feature information.
[0054] The target detection method provided in the embodiment of the present application can adaptively focus on different parts of the features of image information and text information, and enhance the extracted image features and text features to improve the comprehensiveness of feature extraction for image information and text information, and make the key features in the image information and text information more prominent, and then fuse the image features and text features to break the barriers between different modalities so that the two complement each other's feature information, thereby making the fused features more comprehensive and rich, so as to improve the ability to understand complex targets, and then further optimize the fused features through the fusion feature processing model to enhance the robustness of the fused features, thereby providing high-quality feature input for the subsequent positioning calculation processing steps of the target to be detected, so as to improve the accuracy and reliability of target detection.
[0055] Figure 3The following is a flowchart of the target detection method provided in the third embodiment of the present application, which differs from the second embodiment in that:
[0056] The image feature intermediate variable information includes image query feature information, image key feature information and image value feature information;
[0057] The text feature intermediate variable information includes text query feature information, text key feature information and text value feature information;
[0058] The step S203 specifically includes:
[0059] Step S301, calculating image-text cross-attention feature information based on the image query feature information, text key feature information and text value feature information.
[0060] In this embodiment, the text key feature information may be first transposed, and then the image query feature information and the transposed text key feature information are dot-multiplied, and then the dot-multiplication result is scaled, and then the result after scaling is dot-multiplied with the text value feature information, and the result of the last dot-multiplication is normalized by a softmax function, and the processed value is output as the image-text cross-attention feature information. The image-text cross-attention feature information can be used to represent the cross-attention weight information mapped from the image to the text.
[0061] Step S302, calculating text-image cross-attention feature information based on the text query feature information, image key feature information and image value feature information.
[0062] In this embodiment, the image key feature information may be first transposed, and then the text query feature information and the transposed image key feature information are dot-multiplied, and the dot-multiplication result is scaled, and then the result after scaling is dot-multiplied with the image value feature information, and the result of the last dot-multiplication is normalized by a softmax function, and the processed value is output as the text-image cross-attention feature information. The text-image cross-attention feature information can be used to represent the cross-attention weight information mapped from text to image.
[0063] The target detection method provided in the embodiment of the present application can integrate image information into text information by calculating the cross-attention feature information of images and texts, enrich the text description of image targets, and enable the computing network to combine the details of the image when detecting targets based on text descriptions, thereby improving the accuracy of detection. By calculating the cross-attention feature information of images and texts, the text semantics can also be integrated into the image features, helping the computing network to understand the meaning of the target when analyzing the image, more accurately focus on the target to be detected, and avoid being disturbed by irrelevant information, thereby achieving deep interaction between image and text information, allowing the computing network to obtain more comprehensive and accurate feature information, thereby improving the performance of target detection and effectively responding to complex and diverse target detection tasks.
[0064] Figure 4 The flowchart of the target detection method provided in the fourth embodiment of the present application is shown. The difference between the fourth embodiment and the third embodiment is that the step S301 specifically includes:
[0065] Step S401, calculating image-text cross feature intermediate variable information according to the image query feature information and text key feature information.
[0066] In this embodiment, the text key feature information may be transposed first, and then the image query feature information is multiplied by the transposed text key feature information, and the multiplication result is output as the image-text cross feature intermediate variable information.
[0067] Step S402, scaling the image-text cross feature intermediate variable information to calculate the image-text cross feature scaled variable information.
[0068] In this embodiment, the dimension information of the text key feature information may be extracted first, and then the value of the dimension information may be squared, and then the reciprocal of the square root result may be calculated, and the reciprocal may be multiplied with the intermediate variable information of the image-text cross feature, and the multiplication result may be used as the image-text cross feature scaling variable information.
[0069] Step S403, calculating the image-text cross-attention feature information according to the image-text cross-feature scaling variable information and the text value feature information.
[0070] In this embodiment, the image-text cross-feature scaling variable information is first multiplied by the text value feature information, and then the multiplication result is used as the independent variable of the softmax function, and the output function value is used as the image-text cross-attention feature information, which is used as the fused feature information for further calculation operations in subsequent calculations.
[0071] The target detection method provided in the embodiment of the present application can integrate the detailed information of the image into the text, realize the information transmission and fusion from image to text, because the text description is usually more abstract during target detection, the introduction of image details can enrich the text features, make the text description of the target more specific and accurate, and help the computing network to more accurately understand the image features of the target, thereby enhancing the model's ability to locate the target. After the position information in the image is integrated into the text, the computing network can better determine the position of the target to be detected in the image when performing target detection based on the text information, reduce positioning deviation, and thus improve the accuracy and reliability of target detection.
[0072] Figure 5 The flowchart of the target detection method provided in the fifth embodiment of the present application is shown. The difference between the fifth embodiment and the third embodiment is that the step S302 specifically includes:
[0073] Step S501, calculating text-image cross feature intermediate variable information according to the text query feature information and image key feature information.
[0074] In this embodiment, the image key feature information may be transposed first, and then the text query feature information is multiplied by the transposed image key feature information, and the multiplication result is output as the text-image cross feature intermediate variable information.
[0075] Step S502: scaling the text-image cross feature intermediate variable information to calculate the text-image cross feature scaled variable information.
[0076] In this embodiment, the dimension information of the image key feature information can be extracted first, and then the value of the dimensional information is squared, and then the reciprocal of the square root result is calculated, and the reciprocal is multiplied with the intermediate variable information of the text-image cross feature, and the multiplication result is used as the text-image cross feature scaling variable information.
[0077] Step S503, calculating the text-image cross-attention feature information according to the text-image cross-feature scaling variable information and the image value feature information.
[0078] In this embodiment, the text-image cross-feature scaling variable information is first multiplied by the image value feature information, and then the multiplication result is used as the independent variable of the softmax function, and the output function value is used as the text-image cross-attention feature information, which is used as the fused feature information for further calculation operations in subsequent calculations.
[0079] The target detection method provided in the embodiment of the present application can integrate the rich semantic information in the text information into the image features, provide semantic guidance for image understanding, and enable the calculation process to distinguish whether the feature information is related to the target to be detected, so as to avoid being disturbed by the complex background of the image information during the target detection process, thereby improving the accuracy and specificity of the target detection, and in subsequent calculations, quickly searching for targets that meet the text requirements in the image information.
[0080] Figure 6 The flowchart of the target detection method provided in the sixth embodiment of the present application is shown, which differs from the first embodiment described above in that:
[0081] The preset feature decoding model includes a preset image feature decoding model and a preset image-text cross feature decoding model;
[0082] The step S104 specifically includes:
[0083] Step S601, obtaining image decoding features according to the target image feature information and a preset image feature decoding model.
[0084] In this embodiment, the preset image feature decoding model may be set based on the self-attention mechanism, and the target image feature information may be used as input data of the image feature decoding model, and the image feature decoding model may be used to calculate and output image decoding features.
[0085] Step S602, obtaining image-text cross-feature decoding information according to the image decoding features, target text feature information and a preset image-text cross-feature decoding model.
[0086] In this embodiment, the preset text-image cross feature decoding model can be set based on the cross attention mechanism. The target text feature information and the image decoding feature can be used as input data of the text-image cross feature decoding model, and the text-image cross feature decoding model is used to calculate and output text-image cross feature decoding information.
[0087] Step S603, obtaining target image information according to the image-text cross feature decoding information and a preset feature parsing model.
[0088] In this embodiment, the preset feature analysis model may be a feedforward neural network. The image-text cross feature decoding information may be used as input data of the feature analysis model, and the target image information is calculated and output through the feature analysis model. The target image information may refer to the position information of the target to be detected in the image information.
[0089] The target detection method provided in the embodiment of the present application receives the fused and enhanced text and image features through an image feature decoding model, and further processes and optimizes the image features on the basis of incorporating text guidance information into the fused image features, so that the key features in the image features are more prominent, and then the text and image features are deeply interacted through the image-text cross-feature decoding model, and the position information of the target to be detected in the image information is output through the feature parsing model, thereby improving the positioning accuracy and effectiveness of the target to be detected, so as to achieve more efficient and accurate satisfaction of customers' target detection needs in the SaaS platform field.
[0090] Corresponding to the method of the above embodiment, Figure 7 A structural block diagram of a target detection device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown. Figure 7 The exemplary target detection device may be an execution subject of the target detection method provided in the aforementioned first embodiment.
[0091] Reference Figure 7 , the target detection device comprises:
[0092] The target information acquisition module 710 is used to acquire the image information and text information of the target to be detected;
[0093] The initial feature information generating module 720 is used to obtain the initial image feature information and the initial text feature information according to the target image information to be detected, the target text information to be detected, the preset image feature extraction model and the preset text feature extraction model;
[0094] A target feature information generating module 730 is used to obtain target image feature information and target text feature information according to the initial image feature information, the initial text feature information, a preset feature enhancement vector and a preset fusion feature processing model; and
[0095] The target image generation module 740 is used to obtain the target image information according to the target image feature information, the target text feature information, a preset feature decoding model and a preset feature parsing model.
[0096] The process of each module in the target detection device provided in the embodiment of the present application realizing its own function can be specifically referred to the aforementioned Figure 1 The description of the first embodiment is not repeated here.
[0097] The target feature information generating module 730 includes:
[0098] An image feature intermediate variable information calculation unit, used to calculate the image feature intermediate variable information according to the initial image feature information, a preset image query vector, a preset image key vector and a preset image value vector;
[0099] A text feature intermediate variable information calculation unit, used to calculate the text feature intermediate variable information according to the initial text feature information, a preset text query vector, a preset text key vector and a preset text value vector;
[0100] A cross-attention feature calculation unit, used to calculate image-text cross-attention feature information and text-image cross-attention feature information according to the image feature intermediate variable information and the text feature intermediate variable information; and
[0101] The target feature information generating unit is used to obtain the target image feature information and the target text feature information according to the image-text cross-attention feature information, the text-image cross-attention feature information and a preset fusion feature processing model.
[0102] The process of each unit in the target feature information generation module 730 provided in the embodiment of the present application realizing its respective function can be specifically referred to in the aforementioned Figure 2 The description of the second embodiment is not repeated here.
[0103] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0104] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0105] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0106] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0107] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish descriptions, and cannot be understood as indicating or suggesting relative importance. It should also be understood that although the terms "first", "second", etc. are used to describe various elements in some embodiments of the present application in the text, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, the first table can be named as the second table, and similarly, the second table can be named as the first table without departing from the scope of the various described embodiments. The first table and the second table are both tables, but they are not the same table.
[0108] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0109] The target detection method provided in the embodiments of the present application can be applied to terminal devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), etc. The embodiments of the present application do not impose any restrictions on the specific types of terminal devices.
[0110] For example, the terminal device can be a station (STAION, ST) in a WLAN, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, a vehicle-mounted device, a vehicle networking terminal, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, a wireless modem card, a TV set top box (STB), a customer premises equipment (CPE) and / or other devices for communicating on a wireless system and a next-generation communication system, such as a mobile terminal in a 5G network or a mobile terminal in a future evolved Public Land Mobile Network (PLMN) network, etc.
[0111] As an example but not limitation, when the terminal device is a wearable device, the wearable device can also be a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as glasses, gloves, watches, clothing and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not just hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are fully functional, large in size, and can achieve complete or partial functions without relying on smartphones, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.
[0112] Figure 8 Schematic diagram of the structure of a terminal device provided by an embodiment of the present application. Figure 8 As shown, the terminal device 8 of this embodiment includes: at least one processor 80 ( Figure 8 Only one is shown in the figure), a memory 81, in which a computer program 82 that can be run on the processor 80 is stored. When the processor 80 executes the computer program 82, the steps in the above-mentioned target detection method embodiments are implemented, such as Figure 1 Alternatively, when the processor 80 executes the computer program 82, the functions of each module / unit in the above-mentioned device embodiments are realized, for example, Figure 7Functions of modules 710 to 740 are shown.
[0113] The terminal device 8 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The terminal device may include, but is not limited to, a processor 80 and a memory 81. Those skilled in the art will appreciate that Figure 8 It is only an example of the terminal device 8 and does not constitute a limitation of the terminal device 8. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device may also include an input sending device, a network access device, a bus, etc.
[0114] The processor 80 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0115] In some embodiments, the memory 81 may be an internal storage unit of the terminal device 8, such as a hard disk or memory of the terminal device 8. The memory 81 may also be an external storage device of the terminal device 8, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 8. Further, the memory 81 may also include both an internal storage unit and an external storage device of the terminal device 8. The memory 81 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory 81 may also be used to temporarily store data that has been sent or is to be sent.
[0116] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0117] An embodiment of the present application also provides a terminal device, which includes at least one memory, at least one processor, and a computer program stored in the at least one memory and executable on the at least one processor, wherein when the processor executes the computer program, the terminal device implements the steps in any of the above-mentioned method embodiments.
[0118] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0119] An embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0120] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0121] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0122] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0123] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0124] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A target detection method, characterized in that: include: Obtaining image information and text information of the target to be detected; Obtaining initial image feature information and initial text feature information according to the target image information to be detected, the target text information to be detected, a preset image feature extraction model, and a preset text feature extraction model; Obtaining target image feature information and target text feature information according to the initial image feature information, the initial text feature information, a preset feature enhancement vector, and a preset fusion feature processing model; The target image information is obtained according to the target image feature information, the target text feature information, a preset feature decoding model and a preset feature parsing model.
2. The target detection method according to claim 1, characterized in that: The preset feature enhancement vector includes a preset image query vector, a preset image key vector, a preset image value vector, a preset text query vector, a preset text key vector and a preset text value vector; The step of obtaining target image feature information and target text feature information according to the initial image feature information, the initial text feature information, the preset feature enhancement vector and the preset fusion feature processing model specifically includes: Calculating image feature intermediate variable information according to the initial image feature information, a preset image query vector, a preset image key vector, and a preset image value vector; Calculating text feature intermediate variable information according to the initial text feature information, a preset text query vector, a preset text key vector, and a preset text value vector; Calculating image-text cross-attention feature information and text-image cross-attention feature information according to the image feature intermediate variable information and the text feature intermediate variable information; According to the image-text cross-attention feature information, the text-image cross-attention feature information and the preset fusion feature processing model, the target image feature information and the target text feature information are obtained.
3. The target detection method according to claim 2, characterized in that: The image feature intermediate variable information includes image query feature information, image key feature information and image value feature information; The text feature intermediate variable information includes text query feature information, text key feature information and text value feature information; The step of calculating the image-text cross-attention feature information and the text-image cross-attention feature information according to the image feature intermediate variable information and the text feature intermediate variable information specifically includes: Calculating image-text cross-attention feature information based on the image query feature information, text key feature information, and text value feature information; According to the text query feature information, the image key feature information and the image value feature information, the text-image cross-attention feature information is calculated.
4. The target detection method according to claim 3, characterized in that: The step of calculating the image-text cross attention feature information according to the image query feature information, the text key feature information and the text value feature information specifically includes: Calculating image-text cross-feature intermediate variable information based on the image query feature information and the text key feature information; Scaling the intermediate variable information of the image-text cross feature to calculate the scaled variable information of the image-text cross feature; The image-text cross-attention feature information is calculated based on the image-text cross-feature scaling variable information and the text value feature information.
5. The target detection method according to claim 3, characterized in that: The step of calculating the text-image cross attention feature information according to the text query feature information, the image key feature information and the image value feature information specifically includes: Calculating text-image cross-feature intermediate variable information based on the text query feature information and the image key feature information; Scaling the intermediate variable information of the text-image cross feature to calculate the scaled variable information of the text-image cross feature; The text-image cross-feature scaling variable information and the image value feature information are used to calculate the text-image cross-attention feature information.
6. The target detection method according to claim 1, characterized in that: The preset feature decoding model includes a preset image feature decoding model and a preset image-text cross feature decoding model; The step of obtaining the target image information according to the target image feature information, the target text feature information, a preset feature decoding model and a preset feature parsing model specifically includes: Obtaining image decoding features according to the target image feature information and a preset image feature decoding model; Obtaining image-text cross-feature decoding information according to the image decoding features, target text feature information, and a preset image-text cross-feature decoding model; The target image information is obtained according to the image-text cross-feature decoding information and a preset feature parsing model.
7. A target detection device, characterized in that: include: The target information acquisition module is used to acquire the image information and text information of the target to be detected; An initial feature information generation module, used to obtain initial image feature information and initial text feature information according to the target image information to be detected, the target text information to be detected, a preset image feature extraction model and a preset text feature extraction model; A target feature information generation module, used to obtain target image feature information and target text feature information according to the initial image feature information, the initial text feature information, a preset feature enhancement vector and a preset fusion feature processing model; as well as The target image generation module is used to obtain the target image information according to the target image feature information, the target text feature information, a preset feature decoding model and a preset feature parsing model.
8. The target detection device according to claim 7, characterized in that: The target feature information generation module includes: An image feature intermediate variable information calculation unit, used to calculate the image feature intermediate variable information according to the initial image feature information, a preset image query vector, a preset image key vector and a preset image value vector; A text feature intermediate variable information calculation unit, used to calculate the text feature intermediate variable information according to the initial text feature information, a preset text query vector, a preset text key vector and a preset text value vector; A cross-attention feature calculation unit, used to calculate image-text cross-attention feature information and text-image cross-attention feature information according to the image feature intermediate variable information and the text feature intermediate variable information; and The target feature information generating unit is used to obtain the target image feature information and the target text feature information according to the image-text cross-attention feature information, the text-image cross-attention feature information and a preset fusion feature processing model.
9. A terminal device, characterized in that: The terminal device includes a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Video generation method and device, electronic equipment, storage medium and product
CN120455800A