Resource search methods, apparatus, equipment, storage media, and computer program products
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-30
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本申请提供了一种资源搜索方法、装置、电子设备、存储介质及计算机程序产品,可以解决相关技术中存在的资源搜索的准确率不高的问题
在上述技术方案中,在获取到用于指示请求搜索的目标资源的搜索数据后,可以从多个资源描述数据中查找到与该搜索数据对应的资源描述数据,进而基于查找到的资源描述数据,从多个候选资源中获得对应的目标资源,其中,资源描述数据用于描述候选资源的关键属性,也就是说,每一个候选资源都可以通过大语言文本模型被描述为对应的至少一个资源描述数据,使得每一个候选资源的整体描述、标识、对象、对象执行的动作、对象的特征、场景、场景的属性以及事件的重要性类型等关键属性均能够被拆分为可检索字段,而不再是依赖单一自然语言描述进行检索,从而能够有效地提升资源搜索的召回率,能够有效地解决相关技术中存在的资源搜索的准确率不高的问题。
Smart Images

Figure CN122570747A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a resource search method, apparatus, electronic device, storage medium, and computer program product. Background Technology
[0002] With the development of computer technology, users can browse an increasingly wider variety of resources, such as videos, pictures, music, web pages, novels, news, and other multimedia resources.
[0003] Taking videos as an example, the commonly used video retrieval method is based on the similarity between the text keywords entered by the user and the text description content corresponding to the video. Once the text description content corresponding to the video does not completely match or the similarity is not high, the relevant video cannot be retrieved. Summary of the Invention
[0004] This application provides a resource search method, apparatus, electronic device, storage medium, and computer program product, which can solve the problem of low accuracy in resource searches in related technologies. The technical solution provided by this application is as follows: According to one aspect of this application, a resource search method includes: acquiring search data; the search data indicating a target resource to be searched; searching for resource description data corresponding to the search data from a plurality of resource description data; the resource description data being generated by calling a large language text model to describe key attributes of candidate resources; the large language text model being a trained machine learning model capable of recognizing a plurality of key attributes of the candidate resources under the guidance of a structured template and describing each key attribute as a corresponding resource description data; and obtaining the target resource from the plurality of candidate resources based on the found resource description data.
[0005] According to one aspect of this application, a resource search apparatus includes: a data acquisition module for acquiring search data; the search data indicating a target resource to be searched; a data search module for searching for resource description data corresponding to the search data from a plurality of resource description data; the resource description data being generated by calling a large language text model to describe key attributes of candidate resources; the large language text model being a trained machine learning model capable of recognizing a plurality of key attributes of the candidate resources under the guidance of a structured template and describing each key attribute as a corresponding resource description data; and a resource retrieval module for obtaining the corresponding target resource from the plurality of candidate resources based on the found resource description data.
[0006] According to one aspect of this application, an electronic device includes at least one processor and at least one memory, wherein the memory stores a computer program that, when executed by the processor, implements the resource search method as described above.
[0007] According to one aspect of this application, a storage medium having a computer program stored thereon, which, when executed by one or more processors, implements the resource search method as described above.
[0008] According to one aspect of this application, a computer program product includes a computer program that, when executed by one or more processors, implements the resource search method as described above.
[0009] The beneficial effects of the above-mentioned technical solution provided in this application are: In the above technical solution, after obtaining the search data indicating the target resource to be searched, the resource description data corresponding to the search data can be found from multiple resource description data. Then, based on the found resource description data, the corresponding target resource can be obtained from multiple candidate resources. The resource description data is used to describe the key attributes of the candidate resources. That is, each candidate resource can be described by a large language text model as at least one corresponding resource description data. This allows the key attributes of each candidate resource, such as its overall description, identifier, object, action performed by the object, characteristics of the object, scene, attributes of the scene, and importance type of the event, to be broken down into searchable fields, instead of relying on a single natural language description for retrieval. This effectively improves the recall rate of resource search and effectively solves the problem of low accuracy of resource search in related technologies. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram based on the implementation environment involved in this application; Figure 2 This is a hardware structure diagram of an electronic device according to an exemplary embodiment; Figure 3 This is a flowchart illustrating a resource search method according to an exemplary embodiment; Figure 4 This is a schematic diagram of a structured template according to an exemplary embodiment; Figure 5 This is a schematic diagram illustrating resource description data output by a large language text model according to an exemplary embodiment; Figure 6 This is a flowchart illustrating another resource search method according to an exemplary embodiment; Figure 7 yes Figure 6 In one embodiment, step 450 is shown in a flowchart. Figure 8 This is a schematic diagram illustrating the specific implementation of a resource search method in an application scenario; Figure 9 This is a structural block diagram of a resource search device according to an exemplary embodiment; Figure 10 This is a structural block diagram of an electronic device according to an exemplary embodiment. Detailed Implementation
[0012] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0013] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this disclosure means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0014] As mentioned earlier, with the development of computer technology, users can browse an increasingly rich variety of resources. Taking video as an example, existing video retrieval methods usually rely on a single natural language description to complete the retrieval. Specifically, a model or annotation system first generates an overall description of the video, and then performs a full-text search or similarity calculation on the video based on the overall description.
[0015] While the methods described above can express the general content of resources, they still have some shortcomings: First, the overall description is usually a long text in natural language with unclear field boundaries, making it difficult for fine-grained keywords input by users to be stably matched with the corresponding semantics in the long text; Second, the overall description is more focused on readability, easily omitting information such as object, action, scene attributes, and the importance type of the event, resulting in insufficient information dimensions available for retrieval, and consequently making it difficult to accurately match the user's true intent; Third, when the retrieval process relies solely on a single natural language description, it is difficult to simultaneously achieve both recall and precision, potentially leading to missed recalls or false recalls; Fourth, if it is desired to supplement the single natural language description with more information dimensions, it often requires the construction of complex rule chains or manual annotation mechanisms, resulting in high modification costs.
[0016] As can be seen from the above, the relevant technologies still have limitations in terms of the accuracy of resource search.
[0017] Therefore, the resource search method provided in this application can effectively improve the recall and accuracy of resource search. Accordingly, the resource search method is applicable to a resource search device, which can be deployed on an electronic device. The electronic device can be a computer device configured with a von Neumann architecture, such as a desktop computer, a laptop computer, or other electronic devices.
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0019] Figure 1 This is a schematic diagram of an implementation environment involved in an image processing method. It should be noted that this implementation environment is merely an example adapted to the present invention and should not be considered as providing any limitation on the scope of the invention.
[0020] The implementation environment may include client 110 and server 130.
[0021] Specifically, client 110 can be run on a client that provides resource browsing functionality. It can be a desktop computer, laptop computer, tablet computer, smartphone, or other electronic device, and there are no restrictions on its use.
[0022] The client provides resource browsing functionality, such as a media player, web browser, text reader, etc. It can be in the form of an application or a webpage. Correspondingly, the user interface for browsing resources on the client can be in the form of a program window or a webpage, without any limitation here.
[0023] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. This server 130 is an electronic device used to provide backend services. For example, in this implementation environment, server 130 can not only provide cloud storage services for candidate resources to client 110, but also provide resource search services to client 110. Of course, depending on actual operational needs, cloud storage services and resource search services can also be performed independently on different server 130s, so that the above two services are completed by different server 130s.
[0024] A communication connection is pre-established between the server 130 and the user terminal 110 via wired or wireless means, and data transmission between the server 130 and the user terminal 110 is realized through this communication connection. The transmitted data includes, but is not limited to, search data, target resources, etc.
[0025] In one application scenario, on client 110, as the client runs, the user can enter search data in the user interface for browsing resources, so as to initiate a request to server 130 to search for the target resource. Through the interaction between client 110 and server 130, client 110 sends a request carrying search data to server 130 to request server 130 to provide resource search services.
[0026] For server 130, after receiving a request from client 110 to search for a target resource, it obtains the search data from the request and calls the resource search service based on the search data to find the resource description data corresponding to the search data from multiple resource description data. Finally, based on the found resource description data, it obtains the corresponding target resource from multiple candidate resources and returns it to client 110 so that the user can browse the target resource based on the user interface.
[0027] Therefore, each candidate resource can be described by a large language text model as at least one corresponding resource description data. This allows key attributes such as the overall description, identifier, object, action performed by the object, characteristics of the object, scene, attributes of the scene, and importance type of the event of each candidate resource to be broken down into searchable fields, instead of relying on a single natural language description for retrieval. This effectively solves the problem of low resource search accuracy in related technologies.
[0028] Of course, in other application scenarios, the above resource search method can also be completed independently by the user terminal 110, which is not a specific limitation here.
[0029] Please see Figure 2 , Figure 2 This is a hardware structure diagram of an electronic device according to an exemplary embodiment. This electronic device is suitable for... Figure 1 The server 130 in the implementation environment is shown.
[0030] It should be noted that this electronic device is merely an example adapted to this application and should not be construed as providing any limitation on the scope of use of this application. Furthermore, this electronic device should not be interpreted as requiring or depending on any specific feature. Figure 2 One or more components of the exemplary electronic device 200 shown.
[0031] The hardware structure of electronic device 200 can vary significantly due to differences in configuration or performance, such as... Figure 2 As shown, the electronic device 200 includes: a power supply 210, an interface 230, at least one memory 250, and at least one central processing unit (CPU) 270.
[0032] Specifically, power supply 210 is used to provide operating voltage for the various hardware devices on electronic device 200.
[0033] Interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices. For example, to perform... Figure 1 The diagram illustrates the interaction between client 110 and server 130 in the implementation environment.
[0034] Of course, in other examples adapted in this application, interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input / output interface 235, and at least one USB interface 237, etc. Figure 2 As shown, this does not constitute a specific limitation.
[0035] The memory 250 serves as a carrier for resource storage and can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored on it include the operating system 251, application programs 253, and data 255, etc., and the storage method can be temporary storage or permanent storage.
[0036] The operating system 251 is used to manage and control the various hardware devices and application programs 253 on the electronic device 200, so as to enable the central processing unit 270 to perform calculations and processing on the massive data 255 in the memory 250. It can be Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0037] Application 253 is a computer program formed by computer-readable instructions based on operating system 251 to perform at least one specific task, and may include at least one module ( Figure 2 (Not shown), each module can contain corresponding computer-readable instructions. For example, the resource search device can be considered as an application 253 deployed on electronic device 200.
[0038] Data 255 can be photos, videos, etc. stored on a disk, or search data, resource description data, etc., stored in memory 250.
[0039] The central processing unit 270 may include one or more processors and is configured to communicate with the memory 250 via at least one communication bus to read computer programs stored in the memory 250, thereby performing operations and processing on massive amounts of data 255 stored in the memory 250. For example, a resource search method may be implemented by the central processing unit 270 reading an application program 253 stored in the memory 250.
[0040] Furthermore, this application can also be implemented through hardware circuits or a combination of hardware circuits and software. Therefore, the implementation of this application is not limited to any specific hardware circuit, software, or combination thereof.
[0041] Please see Figure 3 This application provides a resource search method, which is applicable to electronic devices. For example, the electronic device may be... Figure 1 The server 130 or user terminal 110 in the implementation environment are shown. The hardware structure of this electronic device can be as follows: Figure 2 As shown. This method can also be implemented through interaction between the user terminal 110 and the server terminal 130.
[0042] In the following method embodiments, for ease of description, the execution subject of each step of the method is the server as an example, but this does not constitute a specific limitation.
[0043] like Figure 3 As shown, the method may include the following steps: Step 310: Obtain search data.
[0044] In this context, search data is used to indicate the target resource requested for the search. The target resource refers to the resource the requester hopes to find; for example, the resource could be a video, image, music, webpage, novel, news, etc. The requester refers to any object that initiates the request to search for the target resource. It should be noted that, depending on the type of requester and resource, the resource search method in this embodiment can be applied to different application scenarios. For example, in a smart home scenario, the resource to be searched could be a video, etc., and the requester could be a family member; in a media resource management scenario, the resource to be searched could be various media resources, such as music, webpages, images, etc., and the requester could be a listener, reader, etc.; in a content review scenario, the resource to be searched could be a novel, news, etc., and the requester could be an author, editor, reviewer, etc.
[0045] In some embodiments, the search data can be in the form of text obtained through text input or in the form of audio obtained through voice input.
[0046] The search data retrieved from the server can originate from real-time user input or from user input pre-stored on the server for a historical period. Figure 1 The illustrated implementation environment demonstrates that after the user receives search data via text or voice input, this data can be sent to the server. The server, upon receiving this search data, can process it in real-time or pre-store it for later processing. For example, it can process the search data when the server's CPU is low, or it can process the search data according to staff instructions. Therefore, the resource search in this embodiment can target real-time search data or search data acquired over a historical period; no specific limitation is made here.
[0047] Step 330: Find the resource description data that corresponds to the search data from multiple resource description data.
[0048] First, it should be noted that resource description data refers to data generated by calling a large language text model to describe the key attributes of candidate resources.
[0049] In some embodiments, key attributes refer to a set of attributes that can be used to reflect the semantic content of a candidate resource. For example, key attributes of a candidate resource may be at least one of the following: an overall description of the candidate resource, an identifier of the candidate resource, objects in the candidate resource, actions performed by the objects, characteristics of the objects, the scene of the candidate resource, attributes of the scene, and the importance type of events in the candidate resource.
[0050] For example, a video of "a man in red walking a black dog in a park" has key attributes including but not limited to: the overall description is "a man and a dog walking in a park", the identifier is "a man and a dog walking", the two objects are "man" and "dog", the characteristics of each object are "man-red clothes" and "dog-black clothes", the action performed by each object is "walking", the scene is "park", the event is "walking" and its importance type is "normal".
[0051] In other words, for each candidate resource, the overall description is a complete description of the objects and / or events within the candidate resource; the identifier can also be considered the name of the candidate resource, used to uniquely identify the candidate resource, and is a simple summary of the objects and / or events within the candidate resource; the object can refer to the people, animals, items, or other things appearing in the candidate resource; the feature refers to the descriptive information used to characterize the appearance characteristics (such as color, clothing), morphological characteristics (such as size, posture), and category characteristics (such as pet breed) of the objects in the candidate resource; the scene refers to the environment in which the objects appear and / or the environment in which the events occur in the candidate resource; the scene attributes refer to the descriptive information used to characterize the environmental features, for example, attributes can be indoor, outdoor, daytime, nighttime, rainy day, or crowded, spacious, etc.; the importance type of the event reflects the degree of abnormality of the event in the candidate resource, and this importance type includes, but is not limited to, ordinary, minor, slight, serious, urgent, etc. For example, in a smart home scene, if the event is "falling," then the importance type of the event is "urgent."
[0052] Secondly, it should be noted that the Large Language Text Model is a trained machine learning model capable of identifying multiple key attributes of candidate resources under the guidance of a structured template, and describing each key attribute as corresponding resource description data. Here, the structured template refers to a template with predefined structured output content and output constraint rules. The output content is used to limit the identifiable elements of the Large Language Text Model (which can also be understood as the key attributes of identifiable candidate resources). The output constraint rules are used to limit the set of output fields that the Large Language Text Model can output, the meaning of the output fields, and the output format and structure of the output fields, given the identifiable elements.
[0053] For example, Figure 4 This illustrates a schematic diagram of a structured template in one embodiment. Figure 4In the structured template, the identifiable elements include the overall description, title, objects, actions, characteristics, and importance of events. Therefore, guided by this structured model, the key attributes of candidate resources that the large language text model can identify include the overall description of the candidate resource, the title of the candidate resource, the objects in the candidate resource, the actions performed by the objects, the characteristics of the objects, and the importance of events in the candidate resource.
[0054] Furthermore, such as Figure 4 As shown, “description”, “title”, “objects”, “actions”, “characteristics”, and “importance” can be considered as the set of output fields that the large language text model can output. “description: A brief description of the overall situation in the video” and “title: A short title following the specified guidelines” can be considered as the meanings of the output fields “description” and “title”, respectively. “actions: A JSON array of strings in format [“subject + action”, “subject + action”,...]” can be considered as the output field “action” with an output format of JSON and an output structure of “object + action”.
[0055] In other words, in this embodiment, the large language text model is a machine learning model with text understanding and generation capabilities. Therefore, guided by structured templates, candidate resources can be transformed into at least one corresponding resource description data. Each resource description data can be considered a structured field describing the key attributes of the candidate resource. Correspondingly, the candidate resource can be considered as a set of structured fields. For example, Figure 5 The illustration shows a schematic diagram of the resource description data corresponding to the candidate resources in one embodiment, such as... Figure 5 As shown, in Figure 4 Guided by the structured template shown, the candidate resource is transformed into six corresponding resource description data. Each resource description data corresponds to a different key attribute of the candidate resource, specifically the overall description of the candidate resource, the identifier of the candidate resource, the object in the candidate resource, the action performed by the object, the characteristics of the object, and the importance type of the event in the candidate resource.
[0056] Therefore, based on at least one resource description data corresponding to each candidate resource, the resource description data corresponding to the search data can be found.
[0057] In some embodiments, resource search is implemented based on matching search data with resource description data. This matching search can include exact match, synonym-normalized match, and qualified field exact match, etc. For example, when the search data is "white pet", the resource description data "white pet" is considered an exact match, while the resource description data "white pet" or "white dog" can be considered a synonym-normalized match, and the resource description data "white dog" can also be considered a qualified field exact match.
[0058] In some embodiments, taking a perfect match as an example, the matching search process may include the following steps: performing a matching search in multiple resource description data based on the search data; if resource description data that matches the search data is found, then it is determined that resource description data corresponding to the search data has been found.
[0059] For example, when the search data is "riding a bicycle", if one of the resource description data A1 corresponding to a certain candidate resource A is also "riding a bicycle", then the resource description data that completely matches the search data "riding a bicycle" can be considered to be A1.
[0060] This approach can effectively improve resource search efficiency in precise matching scenarios.
[0061] In some embodiments, resource search is implemented based on a similarity search between search data and resource description data. Specifically, this may include the following steps: performing a similarity search among multiple resource description data based on the search data; if resource description data similar to the search data is found, then it is determined that resource description data corresponding to the search data has been found.
[0062] In some embodiments, the similarity between search data and resource description data can be determined by directly calculating the similarity between the search data and the resource description data. For example, when the search data is "cycling", if one of the resource description data B1 corresponding to a candidate resource B is "riding a bicycle", by calculating the similarity between "cycling" and "riding a bicycle", the resource description data similar to the search data "cycling" can be considered as B1.
[0063] In some embodiments, the similarity between search data and resource description data can be determined based on calculating the similarity between a first vector corresponding to the search data and a second vector corresponding to the resource description data. Specifically, this may include the following steps: based on the first vector corresponding to the search data, performing a vector search among multiple second vectors; if a second vector similar to the first vector is found, then the resource description data corresponding to the searched second vector is determined as the resource description data corresponding to the search data. Here, the first vector is the semantic vector representation of the search data, and the second vector is the semantic vector representation of the resource description data.
[0064] In some embodiments, vectorization can be achieved through one-hot encoding, word embedding-based Word2Vec, or deep learning pre-trained models such as BERT.
[0065] In some embodiments, similarity is used to measure the semantic closeness of two vectors. This similarity can be calculated using cosine similarity, dot product similarity, or other similarity calculation methods, without limitation. In some embodiments, taking similarity-based similarity search as an example, the specific steps may include: vectorizing the search data to obtain a first vector corresponding to the search data; calculating the similarity between each second vector and the first vector; determining the second vectors similar to the first vector based on the calculated similarity; and identifying the resource description data corresponding to the determined second vectors as resource description data similar to the search data, thereby confirming that resource description data corresponding to the search data has been found.
[0066] Using the example above, suppose the first vector corresponding to the search data "cycling" is a1, and the second vector corresponding to the resource description data "cycling" is b1. By calculating the similarity between the first vector a1 and the second vector b1, we can consider the resource description data that is similar to the search data "cycling" to be B1.
[0067] Of course, in other embodiments, it is not limited to determining the second vector with the highest similarity as the second vector similar to the first vector. Alternatively, several second vectors with similarity exceeding a set threshold can be selected first, and then factors such as the event importance type in the candidate resources associated with the resource description data corresponding to the selected second vector can be considered to comprehensively determine the second vector similar to the first vector.
[0068] This approach helps improve resource search capabilities in semantically similar scenarios, enabling resource search to enhance its ability to identify similar semantic content under different expressions.
[0069] In some embodiments, resource search can first perform a matching search based on search data and resource description data. If a matching resource description data is found, step 350 is executed. If no matching resource description data is found, a similarity search is then performed based on the search data and resource description data, and step 350 is executed based on the similar resource description data found.
[0070] This approach can improve the adaptability of resource search to search data with different expressions, and reduce the poor recall rate of resource search caused by expression differences.
[0071] Step 350: Based on the found resource description data, obtain the corresponding target resource from multiple candidate resources.
[0072] The target resource refers to the candidate resource associated with the found resource description data. It can also be understood as the candidate resource that is ultimately returned to the requester.
[0073] In some embodiments, associating candidate resources with resource description data means that there is a correspondence between the candidate resources and the resource description data. Therefore, the server can search among multiple candidate resources for a candidate resource associated with the found resource description data based on this correspondence, and then return the searched candidate resource as the target resource to the requester.
[0074] As in the example above, the resource description data A1 "riding a bicycle" perfectly matches the search data "riding a bicycle". Based on the correspondence between the resource description data A1 and the candidate resource A, the server can search among multiple candidate resources and find the candidate resource A that corresponds to the resource description data A1, and then determine the candidate resource A as the target resource.
[0075] In some embodiments, associating candidate resources with resource description data means that there is a correspondence between the candidate resources and the second vector corresponding to the resource description data. Therefore, the server can search among multiple candidate resources for candidate resources that are associated with the found resource description data, based on the correspondence between the candidate resources and the second vector corresponding to the resource description data, and then return the searched candidate resource as the target resource to the requester.
[0076] As in the example above, similar to the first vector a1 corresponding to the search data "riding a bicycle", the second vector b1 corresponding to the resource description data B1 "riding a bicycle" is similar. Based on the correspondence between the resource description data B1 and the candidate resource B, the server can search for the candidate resource B that corresponds to the resource description data B1 among multiple candidate resources, and then determine the candidate resource B as the target resource.
[0077] In some embodiments, the correspondence can be stored using identifier mapping tables, key-value pairs, etc., so that the server can search for associated candidate resources based on resource description data. For example, in the identifier mapping table, the resource description data "riding a bicycle" corresponds to identifier A. Therefore, the server can locate candidate resource A in a resource library storing a large number of candidate resources based on identifier A, and return candidate resource A to the requester. Alternatively, the resource description data "riding a bicycle" can be stored as a key, and candidate resource A as the value corresponding to the key. Then, based on the key-value pair, candidate resource A can be directly returned to the requester from the resource description data "riding a bicycle".
[0078] Through the above process, each candidate resource can be described by a large language text model as at least one corresponding resource description data. This allows key attributes such as the overall description, identifier, object, action performed by the object, characteristics of the object, scene, attributes of the scene, and importance type of the event of each candidate resource to be broken down into searchable fields, instead of relying on a single natural language description for retrieval. This effectively improves the recall rate of resource search and effectively solves the problem of low accuracy in resource search in related technologies.
[0079] Please see Figure 6 In an exemplary embodiment, prior to step 330, the method may further include the following steps: Step 410: Obtain the resources to be processed.
[0080] Here, the resources to be processed refer to resources that have not yet been processed by the large language text model. In some embodiments, the types of resources to be processed include, but are not limited to, image types, audio types, and text types. For example, image types can be videos (moving images), pictures (still images), etc., audio types can be music, speech, etc., and text types can be web pages, novels, news, etc.
[0081] It should be noted that the source of the resources to be processed can vary depending on the type of resource. Taking video as an example, the resources to be processed can come from surveillance video playback, home camera captures, short video content streams, or media material libraries, etc.
[0082] In some embodiments, the resource to be processed may be video captured in real time by an image acquisition device (such as a camera). In other embodiments, the resource to be processed may also be video captured by an image acquisition device over a historical period, obtained in batches from the resource upload interface on the user's end.
[0083] Step 430: Extract features from the resource to be processed to obtain the multimodal features of the resource to be processed.
[0084] Multimodal features refer to feature representations that reflect the content, temporal sequence information, and other available modal information of the resource to be processed from multiple dimensions. For example, if the resource to be processed is a video, the multimodal features of the video can include visual features and temporal sequence features.
[0085] Taking video as an example of the resource to be processed, the process of extracting multimodal features from video may include the following steps: The first step is to extract visual features from multiple image frames in the video to be processed, thereby obtaining the visual features of the video. During the visual feature extraction process, the temporal dependencies between the image frames are preserved to obtain the temporal features of the video to be processed.
[0086] Among them, visual features reflect the visual semantic information of the video. Temporal features reflect the temporal order information of each image frame in the video.
[0087] In some embodiments, visual feature extraction can be based on a video encoder to convert each image frame in the video into high-dimensional features, i.e., visual features.
[0088] In some embodiments, temporal feature extraction can be achieved by introducing a spatiotemporal attention mechanism on the basis of the visual encoder, so as to capture the correlation (i.e., temporal dependency) between different image frames in the time dimension during the visual feature extraction process, thereby perceiving the continuity between different image frames and avoiding feature loss that would affect the subsequent processing of large language text models.
[0089] The second step is to fuse the visual features and temporal features of the video to be processed using positional encoding to obtain the multimodal features of the video.
[0090] Among them, the positional encoding method refers to the encoding method that adds time sequence identification information (depending on temporal features) to different image frames (depending on visual features) to preserve the temporal order of different image frames during feature fusion.
[0091] In other words, by fusing visual and temporal features, multimodal features not only accurately describe the content of each image frame in the video from a spatial dimension, but also accurately describe the relationship between each image frame in the video from a temporal dimension. This is beneficial for subsequent large language text models to accurately translate the video into corresponding resource description data.
[0092] Of course, in other embodiments, the image frames for feature extraction are not limited to all image frames in the video, but may also be at least a few key frames in the video to reduce processing costs.
[0093] In some embodiments, the keyframe extraction process may include the following steps: first, sampling the video to be processed to obtain multiple image frames in the video; then, extracting at least one keyframe from the multiple image frames. Here, a keyframe refers to an image frame that can represent a change in the image content or the occurrence of a major event in the video.
[0094] It should be noted that sampling processing can be implemented by extracting a set number of image frames from the video to be processed according to a set sampling frequency, for example, setting the sampling method to one frame per second. Keyframe extraction can be implemented based on differences in image content (such as the inter-frame difference method), or based on time domain methods (such as the shot boundary method), or based on transform domain methods (such as the wavelet transform method), or based on semantic / deep learning methods (such as the autoencoder), etc., without limitation here. Taking the inter-frame difference method based on pixel differences as an example, if the pixel difference between adjacent image frames a and b exceeds a set threshold, then image frame b can be regarded as a keyframe, which can be considered as an image frame with significant image changes.
[0095] Therefore, it is possible to retain key semantic information and temporal sequence information in the video while controlling computational costs, thus serving as the data basis for text understanding and generation in large language text models.
[0096] Step 450: Guided by the structured template, the multimodal features of the resource to be processed in the input large language text model are transformed into at least one resource description data of the resource to be processed.
[0097] As mentioned earlier, a structured template refers to a template that predefines structured output content and output constraint rules. The output content is used to limit the identifiable elements of the large language text model (which can also be understood as the key attributes of identifiable candidate resources). The output constraint rules are used to limit the set of output fields that the large language text model can output, the meaning of the output fields, and the output format and structure of the output fields, given the identifiable elements. Correspondingly, each resource description data can be considered as a structured field describing the key attributes of a candidate resource.
[0098] Based on this, guided by the predefined output content and output constraint rules of the structured model, the large language text model actually transforms at least one key attribute of the resource to be processed into at least one structured field, that is, at least one resource description data. Figure 7 A flowchart of the conversion process in one embodiment is shown, such as Figure 7 As shown, the conversion process may include the following steps: Step 451: According to the structured template, the multimodal features of the resource to be processed are input into the large language text model to identify at least one key attribute of the resource to be processed, thereby obtaining at least one key attribute of the resource to be processed.
[0099] In other words, based on the identifiable elements (i.e. key attributes) of the large language text model defined in the structured template, the large language text model will identify the identifiable elements of the resource to be processed and output a set of output fields that conforms to the structured template and contains multiple output fields, as well as the meaning of each output field in the set of output fields.
[0100] In this embodiment, at least one key attribute of the resource to be processed identified by the large language text model is represented by an output field set containing at least one output field and the meaning of each output field in the output field set.
[0101] Please refer back to Figure 4 ,exist Figure 4 In a structured template, identifiable elements include the overall description, title, objects, actions, characteristics, and importance. Taking a video of a man walking his dog as an example, under the guidance of this structured model, the key attributes of the resource identified by the large language text model are actually the set of output fields that the large language text model can output, as well as the meaning of each output field in the set of output fields. These include: the output field description, which represents the overall description of the resource, meaning "A man is walking in a park with his dog"; the output field title, which represents the identifier of the resource, meaning "man walking dog"; the output field objects, which represents the objects in the resource, meaning "person", "dog", "trees", and "path"; the output field actions, which represents the actions performed by the objects, meaning "walking" performed by the object "person" and "moving" performed by the object "dog"; the output field characteristics, which represents the characteristics of the objects, meaning "blue shirt" for the object "person", "brown" for the object "dog", and "green" for the object "trees"; and the output field importance, which represents the importance type of the event, meaning "normal".
[0102] Step 453: In the large language text model, according to the structured description rules indicated by the structured template, the key attributes of the resource to be processed are structured to obtain the resource description data of the resource to be processed.
[0103] Among them, the structured description rule refers to the output constraint rule of the structured template for the predefined output content. It can also be understood as the structured description rule used to limit the output format and output structure of each output field that the large language text model can output.
[0104] In some embodiments, structured processing refers to converting each key attribute of the resource to be processed into at least one corresponding structured field. This structured field refers to a field that conforms to the output format and structure of the structured description rules.
[0105] Continuing with the previous example, in the structured template, the output constraint rule for the `actions` field is "actions: A JSON array of strings in format ["subject + action", "subject + action",...]". Please refer back to... Figure 5 ,exist Figure 5 In the output field actions—the action "walking" performed by the object "person" and the action "moving" performed by the object "dog"—structured processing generates structured fields "a man is walking" and "a dog is moving" in JSON format, which are regarded as resource description data corresponding to the actions performed by the objects.
[0106] This allows resources that are difficult to retrieve directly to be converted into structured fields that are understandable, storable, and searchable, thereby improving the searchability and parsability of the resources.
[0107] It should be noted that any key attribute of the resource to be processed can correspond to one resource description data. For example, if the key attribute is an overall description, identifier, scenario, or event importance type, then the key attribute corresponds to one resource description data. Alternatively, it can correspond to multiple resource description data. For example, if the key attribute is an object, the action performed by the object, the characteristics of the object, or the attributes of the scenario, then the key attribute can correspond to at least one resource description data. No limitation is imposed here.
[0108] Step 470: By establishing a correspondence between the resource to be processed and multiple resource description data, and / or by establishing a correspondence between the resource to be processed and the second vector corresponding to each resource description data, the resource to be processed is transformed into a candidate resource.
[0109] The second vector is obtained by vectorizing the resource description data.
[0110] In other words, after the large language text model outputs multiple resource description data for the resource to be processed, the resource to be processed can be transformed into candidate resources through the establishment of a correspondence. This correspondence can refer to the correspondence between the resource description data and the resource to be processed, or it can refer to the correspondence between the second vector corresponding to the resource description data and the resource to be processed.
[0111] In some embodiments, the correspondence can be stored using identifier mapping tables, key-value pairs, etc., so that the server can subsequently search for related candidate resources based on the resource description data. For example, in the identifier mapping table, the resource description data "riding a bicycle" corresponds to identifier A. Therefore, the server can locate candidate resource A in a resource library storing a massive number of candidate resources based on identifier A, and return candidate resource A to the requester. Alternatively, the resource description data "riding a bicycle" can be stored as a key, and candidate resource A as the value corresponding to the key. Then, based on the key-value pair, candidate resource A can be directly returned to the requester from the resource description data "riding a bicycle".
[0112] By combining the above embodiments, a complete mapping is established from the structured description of the resource to be processed to the resource description data and then to the construction of the corresponding relationship. This enables the addition of the resource to be processed to the candidate resource set, allowing subsequent resource search based on the resource description data to be realized, which is beneficial to improving the accuracy and recall rate of resource search.
[0113] Figure 8 This is a schematic diagram of a resource search method in an application scenario. In this application scenario, the resource search method can be implemented through a resource search system, which includes a user terminal 510 for requesters to input search data to request the search for target resources, and a server cluster.
[0114] The server cluster comprises multiple servers providing different backend services: a resource repository 530, a structured parsing server 550, a vector indexing server 570, and a resource search server 590. The resource repository 530 stores candidate resources such as videos; the structured parsing server 550 uses a large language text model to generate at least one resource description data for each candidate resource based on a structured template; the vector indexing server 570 stores a first vector corresponding to the search data and a second vector corresponding to the resource description data through vectorization processing; and the resource search server 590 searches for target resources based on the search data.
[0115] like Figure 8 As shown, the resource library 530 first requests the structured parsing server 550 to convert the resource to be processed into a candidate resource, and then stores the converted candidate resource.
[0116] Taking video as an example of the resource to be processed, video data can first be obtained from the video source (surveillance, short video, live playback, etc.). Then, the video is sampled (e.g., 1-3 frames per second), and key frames are extracted to reduce processing costs. Specifically, key frame extraction can be performed by storing one current frame per second, converting it to grayscale, and calculating the difference between the current frame and the grayscale image stored the previous second. When the difference is greater than 30 (an adjustable threshold), it is recorded. If more than 5% of the pixels have a difference greater than 30, the current frame is considered a key frame (i.e., a frame with significant image changes).
[0117] Then, a multimodal large model can be used to extract multimodal features, which can specifically include visual features and time-series information. Visual features can be frame-level visual information extracted based on a video encoder, while time-series information is used to preserve temporal dependencies between frames. The obtained multimodal features are then fed into a large language text model for processing. The time-series information is added to the visual features using positional encoding to obtain the final token fed into the large language text model. The large language text model can be an LLM model, which accepts a token as input and outputs a token that can be translated into text using a dictionary.
[0118] By pre-configuring structured templates, such as the Prompt template, multimodal large models can be guided to directly output JSON or table formats that conform to a predefined schema. For example, the Prompt can explicitly require the model to cover all identifiable elements. These identifiable elements include fields such as overall description, title, objects, actions, characteristics, and importance of events. These identifiable elements are mainly used for subsequent search needs, rather than just generating a short summary.
[0119] Then, the structured fields described above are vectorized and the output is stored in a database or search engine (such as Elasticsearch, Milvus+ Vector Search). For the vectorization method, the structured fields can be split: description, title, and importance each generate a text vector; each document in the actions, characters, and objects lists generates a text vector.
[0120] When users search for videos, they can use natural language vector searches or phrase matching searches. If using vector search, for example, if a camera uploads a video of a user riding a bicycle, a multimodal large-scale model can process the video to obtain fields such as description, actions, character, and labels. These fields are then vectorized (using a text embedding model), potentially generating a dozen or more vectors for a single video, which are then stored in a vector database. When a user searches, their search field is also vectorized, and the database is matched against the closest vector. If a match is found, the corresponding video is retrieved. Because each video has multiple vectors of different dimensions, the recall rate can be significantly improved.
[0121] If a matching search method is used, for example, if the user's search terms exactly match the content of fields such as actions, characters, or objects, the video can be retrieved directly without vector comparison, which can greatly improve the user's search efficiency.
[0122] For example, suppose there is a video of "a man in a red shirt walking in a park with his black dog". Guided by a predefined Schema / Prompt structured template, the structured parsing server 550 calls the large language text model to convert the different key attributes of the video into corresponding structured fields (e.g., description, actions, objects, characteristics, title, importance), i.e., resource description data. For example, for this video, the key attribute - object can be transformed into the structured fields "objects" and "dog" in the resource description data; the key attribute - object characteristics can be transformed into the structured fields "red clothes" and "black" in the resource description data; the key attribute - identifier can be transformed into the structured field "person and dog walking" in the resource description data; the key attribute - action can be transformed into the structured fields "person + walking" and "dog + walking" in the resource description data; the key attribute - overall description can be transformed into the structured field "a man in red clothes is walking a black dog in the park" in the resource description data; and the key attribute - event importance type can be transformed into the structured field "importance" in the resource description data. Thus, the video is transformed into multiple resource description data. Simultaneously, the structured parsing server 550 will also verify and correct the transformed resource description data to ensure that this resource description data is in machine-parseable JSON format.
[0123] For resource search server 590, it can retrieve the video based on the above-mentioned multiple resource description data.
[0124] For example, if the search data input by the requester 510 is "a man in a red shirt and his black dog walking in the park", the resource search server 590 can obtain resource description data that matches the search data through matching search, including "a man in a red shirt walking with a black dog in the park", "man", "dog", "red shirt", "black", "park", "person + walking", "dog + walking", and "ordinary". Based on this resource description data, the resource search server 590 can obtain the identifier of the video through candidate resource mapping, and then retrieve the video from the resource library 530 through candidate resource positioning.
[0125] Similarly, if the search data input by the requester 510 is "a person walking a large dog on the road", similar search can obtain resource description data similar to the search data, including "man", "dog", "person + walking", "dog + walking", and "ordinary". Therefore, the resource search server 590 can also obtain the identifier of the video by candidate resource mapping based on these resource description data, and then retrieve the video from the resource library 530 by candidate resource location.
[0126] The vectorization process described above is performed by the vector index server 570.
[0127] In this application scenario, the video is automatically converted into multiple structured fields that can be understood, stored, and retrieved under the guidance of structured templates. This makes the video results no longer dependent on a single natural language text description, but can be jointly retrieved and located based on multiple structured fields, which helps to improve the recall and accuracy of video retrieval.
[0128] Compared with related technologies, the embodiments of this application have at least the following beneficial effects: First, by representing candidate resources as multiple resource description data instead of a single natural language description, information such as the overall description, object, action, object features, scene, scene attributes, and event importance type can be broken down into searchable fields, thereby improving the recall rate of resource retrieval; Second, by combining matching search and vector search, both the needs for precise keyword retrieval and the needs for retrieval of semantically similar expressions can be met, thereby improving search accuracy and retrieval flexibility; Third, by constructing a correspondence between candidate resources, resource description data, and the second vector, the corresponding candidate resources can be quickly traced back after finding the relevant description, thereby improving resource recall efficiency; Fourth, in video scenarios, by extracting keyframes, fusing multimodal features, and generating structured template constraints, rich semantic information can be retained while controlling computational costs, thereby improving the quality of video search; Fifth, this application can add structured parsing and indexing capabilities to the existing search architecture without completely reconstructing the search system, thus facilitating engineering implementation.
[0129] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0130] The following are embodiments of the apparatus described in this application, which can be used to execute the resource search method involved in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the method embodiments of the resource search method involved in this application.
[0131] Please see Figure 9 This application provides a resource search device 900, including but not limited to: a data acquisition module 910, a data search module 930, and a resource retrieval module 950.
[0132] Among them, the data acquisition module 910 is used to acquire search data; the search data is used to indicate the target resource requested for the search; The data search module 930 is used to find resource description data corresponding to the search data from multiple resource description data; the resource description data is generated by calling the large language text model and is used to describe the key attributes of the candidate resources; the large language text model is a machine learning model that has been trained and has the ability to identify multiple key attributes of the candidate resources under the guidance of a structured template and describe each key attribute as the corresponding resource description data. The resource retrieval module 950 is used to obtain the corresponding target resource from multiple candidate resources based on the found resource description data.
[0133] In some embodiments, the data search module 930 is further configured to perform a matching search among multiple resource description data based on the search data; if resource description data matching the search data is found, it is determined that resource description data corresponding to the search data has been found.
[0134] In some embodiments, the data search module 930 is further configured to perform a vector search among a plurality of second vectors based on a first vector corresponding to the search data; the second vector is obtained by vectorizing the resource description data; if a second vector similar to the first vector is found, the resource description data corresponding to the searched second vector is determined as the resource description data corresponding to the search data.
[0135] In some embodiments, the data search module 930 is further configured to perform vectorization processing on the search data to obtain the first vector corresponding to the search data; calculate the similarity between each second vector and the first vector respectively; and determine the second vector similar to the first vector based on the calculated similarity.
[0136] In some embodiments, the resource retrieval module 950 is further configured to search for candidate resources associated with the found resource description data among a plurality of candidate resources based on the correspondence between the candidate resources and the resource description data, and / or based on the correspondence between the candidate resources and the second vector corresponding to the resource description data; the second vector is obtained by vectorizing the resource description data; and the searched candidate resources are determined as the target resources.
[0137] In some embodiments, the resource search device further includes a resource acquisition module, configured to acquire a resource to be processed; extract features from the resource to be processed to obtain multimodal features of the resource to be processed; under the guidance of the structured template, convert the multimodal features of the resource to be processed input into the large language text model into at least one resource description data of the resource to be processed; and convert the resource to be processed into the candidate resource by constructing a correspondence between the resource to be processed and multiple resource description data of the resource to be processed, and / or constructing a correspondence between the resource to be processed and the second vectors corresponding to each resource description data.
[0138] In some embodiments, the resource to be processed includes a video to be processed; the resource acquisition module is further configured to sample the video to be processed to obtain multiple image frames in the video to be processed; and to extract at least one key frame of the video to be processed from the multiple image frames, so that the feature extraction is based on at least one key frame of the video to be processed.
[0139] In some embodiments, the resource to be processed includes a video to be processed; the resource search device further includes a feature extraction module, which is used to extract visual features from multiple image frames in the video to be processed to obtain the visual features of the video to be processed, and retain the temporal dependency between each image frame during the visual feature extraction process to obtain the temporal features of the video to be processed; and fuse the visual features and the temporal features of the video to be processed through a position encoding method to obtain the multimodal features of the video to be processed.
[0140] In some embodiments, the resource acquisition module is further configured to input the multimodal features of the resource to be processed into the large language text model according to at least one key attribute that can be identified by the large language text model indicated by the structured template, identify the key attributes of the resource to be processed, and obtain at least one key attribute of the resource to be processed; in the large language text model, according to the structured description rules indicated by the structured template, perform structured processing on each key attribute of the resource to be processed to obtain each resource description data of the resource to be processed.
[0141] In some embodiments, the key attributes include at least one of the following: an overall description of the candidate resource, an identifier of the candidate resource, an object in the candidate resource, an action performed by the object, characteristics of the object, a scene of the candidate resource, attributes of the scene, and importance type of the event in the candidate resource.
[0142] It should be noted that the resource search device provided in the above embodiments is only illustrated by the division of the above functional modules when performing resource searches. In actual applications, the above functions can be assigned to different functional modules as needed. That is, the internal structure of the resource search device will be divided into different functional modules to complete all or part of the functions described above.
[0143] Furthermore, the resource search device and resource search method embodiments provided in the above embodiments belong to the same concept, and the specific way in which each module performs operations has been described in detail in the method embodiments, and will not be repeated here.
[0144] Please see Figure 10 This application provides an electronic device 4000, which may include a server, etc.
[0145] exist Figure 10 In this context, the electronic device 4000 includes at least one processor 4001 and at least one memory 4003.
[0146] Data interaction between the processor 4001 and the memory 4003 can be achieved through at least one communication bus 4002. This communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not indicate that there is only one bus or one type of bus.
[0147] Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.
[0148] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0149] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing computer programs having instruction or data structure forms and accessible by the electronic device 4000, but not limited to these.
[0150] The memory 4003 stores a computer program, and the processor 4001 can read the computer program stored in the memory 4003 through the communication bus 4002.
[0151] The computer program is executed by one or more processors 4001 to implement the resource search methods in the above embodiments.
[0152] Furthermore, this application provides a storage medium storing a computer program, which is executed by one or more processors to implement the resource search method described above.
[0153] This application provides a computer program product, including a computer program that is executed by one or more processors to implement the resource search method described above.
[0154] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A resource search method, characterized in that, include: Obtain search data; the search data is used to indicate the target resource for which a search is requested; The resource description data corresponding to the search data is found from multiple resource description data; the resource description data is generated by calling a large language text model to describe the key attributes of the candidate resources; the large language text model is a trained machine learning model that has the ability to identify multiple key attributes of the candidate resources under the guidance of a structured template and describe each key attribute as the corresponding resource description data. Based on the found resource description data, the corresponding target resource is obtained from multiple candidate resources.
2. The method as described in claim 1, characterized in that, The step of searching for the resource description data corresponding to the search data from multiple resource description data includes: Based on the search data, a matching search is performed among multiple resource description data; If resource description data that matches the search data is found, then it is determined that resource description data corresponding to the search data has been found.
3. The method as described in claim 1, characterized in that, The step of searching for the resource description data corresponding to the search data from multiple resource description data includes: Based on the first vector corresponding to the search data, a vector search is performed among multiple second vectors; the second vector is obtained by vectorizing the resource description data. If a second vector similar to the first vector is found, the resource description data corresponding to the second vector is determined as the resource description data corresponding to the search data.
4. The method as described in claim 3, characterized in that, The vector search based on the first vector corresponding to the search data, among multiple second vectors, includes: The search data is vectorized to obtain the first vector corresponding to the search data; Calculate the similarity between each of the second vectors and the first vector; Based on the calculated similarity, a second vector similar to the first vector is determined.
5. The method according to any one of claims 1 to 4, characterized in that, The step of obtaining the corresponding target resource from multiple candidate resources based on the found resource description data includes: Based on the correspondence between the candidate resources and the resource description data, and / or based on the correspondence between the candidate resources and the second vector corresponding to the resource description data, the candidate resources associated with the found resource description data are searched among a plurality of candidate resources; the second vector is obtained by vectorizing the resource description data; The candidate resources found in the search are identified as the target resources.
6. The method according to any one of claims 1 to 4, characterized in that, Before searching for the resource description data corresponding to the search data from multiple resource description data, the method further includes: Acquire resources to be processed; Feature extraction is performed on the resource to be processed to obtain its multimodal features; Guided by the structured template, the multimodal features of the resource to be processed input into the large language text model are transformed into at least one resource description data of the resource to be processed; By establishing a correspondence between the resource to be processed and multiple resource description data of the resource to be processed, and / or establishing a correspondence between the resource to be processed and the second vector corresponding to each of the resource description data, the resource to be processed is transformed into the candidate resource.
7. The method as described in claim 6, characterized in that, The resources to be processed include videos to be processed; After acquiring the resource to be processed, the method further includes: The video to be processed is sampled to obtain multiple image frames in the video to be processed; At least one keyframe of the video to be processed is extracted from the plurality of image frames, such that the feature extraction is performed based on at least one keyframe of the video to be processed.
8. The method as described in claim 6, characterized in that, The resources to be processed include videos to be processed; The step of extracting features from the resource to be processed to obtain the multimodal features of the resource to be processed includes: Visual features are extracted from multiple image frames in the video to be processed to obtain the visual features of the video to be processed. The temporal dependency between each image frame is preserved during the visual feature extraction process to obtain the temporal features of the video to be processed. By using positional encoding, the visual features of the video to be processed are fused with the temporal features to obtain the multimodal features of the video to be processed.
9. The method as described in claim 6, characterized in that, Guided by the structured template, the process of converting the multimodal features of the resource to be processed into at least one resource description data of the resource to be processed, including: According to at least one of the key attributes that the large language text model can identify as indicated by the structured template, the multimodal features of the resource to be processed are input into the large language text model to identify the key attributes of the resource to be processed, thereby obtaining at least one of the key attributes of the resource to be processed. In the large language text model, according to the structured description rules indicated by the structured template, the key attributes of the resource to be processed are structured to obtain the resource description data of the resource to be processed.
10. The method according to any one of claims 1 to 4, characterized in that, The key attributes include at least one of the following: an overall description of the candidate resource, the identifier of the candidate resource, the object in the candidate resource, the action performed by the object, the characteristics of the object, the scenario of the candidate resource, the attributes of the scenario, and the importance type of the event in the candidate resource.
11. A resource search device, characterized in that, include: The data acquisition module is used to acquire search data; the search data is used to indicate the target resource for which a search is requested. A data search module is used to find the resource description data corresponding to the search data from multiple resource description data; the resource description data is generated by calling a large language text model and is used to describe the key attributes of candidate resources; the large language text model is a trained machine learning model that has the ability to identify multiple key attributes of the candidate resources under the guidance of a structured template and describe each key attribute as the corresponding resource description data; The resource retrieval module is used to obtain the corresponding target resource from multiple candidate resources based on the found resource description data.
12. An electronic device comprising at least one processor and at least one memory, wherein, The memory stores a computer program, characterized in that, when the computer program is executed by the processor, it implements the resource search method as described in any one of claims 1 to 10.
13. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by one or more processors, it implements the resource search method as described in any one of claims 1 to 10.
14. A computer program product comprising a computer program, characterized in that, When the computer program is executed by one or more processors, it implements the resource search method as described in any one of claims 1 to 10.