Manuscript content retrieval method and device
By building a cross-modal target content detection model, a fine-grained semantic understanding of video, images and text manuscripts is achieved, and the problem of poor retrieval of manuscript content in the existing technology is solved, and the compliance review capabilities of the content creation platform are improved.
Patent Information
- Application Number
- CN202510413208.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, the graphic encoding model cannot effectively capture the detailed information in videos, images and text, resulting in poor retrieval of manuscript content and cannot meet the compliance review needs of the content creation platform.
A cross-modal target content detection model is constructed, the characteristic information of the manuscript is obtained through multi-dimensional semantic analysis, and compliance detection is performed, and the prediction probability is output to determine the manuscript search results.
It improves the efficiency and accuracy of manuscript content retrieval, can automatically recall content containing target semantics, improves the recall ability of users' malicious content and new categories of samples, and supports the search of multiple modal manuscripts.
Smart Images

Figure CN120407891A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of Internet technologies, and in particular, to a method, device, computer device, computer-readable storage medium, and computer program product for retrieving manuscript content. Background Art
[0002] In the prior art, generally, based on video / image-text Pair data, a unified text and image encoding model is trained to obtain a trained text and image encoding model to support functions such as text-to-image, image-to-text, text-to-text, and image-to-image. The text and image encoding model trained by the above solution can only support a relatively coarse-grained semantic understanding of the input data, and many detailed information cannot be covered, such as some content information hidden in a nine-square grid image, and when the input information in the query contains too many details, it will also be ignored.
[0003] Currently, many content creation platforms need to first review the content uploaded by users to conduct compliance inspections on the content. The model capabilities for supporting content review services include various vertical classification models, image retrieval models, text matching, large language models, and multi-modal classification models, etc. However, there is no general model that can take into account all the detailed information contained in the video, image, and text of the manuscript, so it is impossible to use these models to recall the target content in the target manuscript, or the recall effect is not good.
[0004] It should be noted that the above content is not necessarily prior art and is not used to limit the patent protection scope of the present application. Summary of the Invention
[0005] Embodiments of the present application provide a method, device, computer device, computer-readable storage medium, and computer program product for retrieving manuscript content to solve or alleviate one or more of the above technical problems.
[0006] One aspect of the embodiments of the present application provides a method for retrieving manuscript content, the method comprising: obtaining a manuscript to be processed and a target content detection model; wherein, the manuscript to be processed is any one of text, image, and video, and the target content detection model is used to support cross-modal target semantic retrieval; performing semantic analysis on the manuscript to be processed from multiple preset dimensions through the target content detection model to obtain feature information of multiple dimensions; performing compliance detection according to the feature information of multiple dimensions to obtain a prediction probability, and outputting a manuscript retrieval result according to the prediction probability.
[0007] Optionally, the method further comprises: When the predicted probability is greater than a preset threshold, the manuscript search result is output as the manuscript to be processed failing the search.
[0008] Optionally, the method further includes: In the case that the manuscript retrieval result belongs to a preset type, a secondary retrieval is performed on the manuscript to be processed through a preset retrieval channel to obtain a target manuscript retrieval result.
[0009] Optionally, before the step of obtaining the manuscript to be processed and the target content detection model, the method further includes: Obtain positive and negative sample data of multiple modalities; Performing feature extraction on the positive sample data of the multiple modalities to obtain first feature information, and performing feature extraction on the negative sample data of the multiple modalities to obtain second feature information; wherein the first feature information and the second feature information belong to a first granularity; An initial content detection model is trained based on the first feature information and the second feature information.
[0010] Optionally, the method further includes: Semantically analyzing the positive sample data of the multiple modalities to obtain third feature information of multiple dimensions, and semantically analyzing the negative sample data of the multiple modalities to obtain fourth feature information of multiple dimensions; wherein the third feature information and the fourth feature information belong to the second granularity; Converting the third feature information and the fourth feature information into a preset vector space; The initial content detection model is trained based on the converted third feature information and fourth feature information to obtain the target content detection model.
[0011] Optionally, the training of the initial content detection model according to the converted third feature information and fourth feature information to obtain the target content detection model includes: Determining the similarity between the converted third feature information and the fourth feature information; Determining multiple model losses based on the similarities according to multiple preset loss functions; Parameters of the initial content detection model are optimized according to the multiple model losses to obtain the target content detection model.
[0012] Optionally, the third feature information and the fourth feature information respectively include: feature information of the detection head dimension, feature information of the classification head dimension and feature information of the vector head dimension; the feature information of the detection head dimension and the feature information of the classification head dimension are used to mine the semantic understanding capability of the second granularity, and the feature information of the vector head dimension is used to align the image and text recall capability.
[0013] Another aspect of the embodiments of the present application provides a manuscript content retrieval device, which includes: A to-be-processed manuscript acquisition module, configured to acquire a to-be-processed manuscript and a target content detection model; wherein, the to-be-processed manuscript is any one of text, image, and video, and the target content detection model is used to support cross-modal target semantic retrieval; A semantic analysis module, configured to perform semantic analysis on the to-be-processed manuscript from multiple preset dimensions through the target content detection model to obtain feature information of multiple dimensions; A compliance detection module, configured to perform compliance detection according to the feature information of multiple dimensions to obtain a prediction probability, and output a manuscript retrieval result according to the prediction probability.
[0014] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the manuscript content retrieval method as described above.
[0015] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the manuscript content retrieval method as described above is implemented.
[0016] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the manuscript content retrieval method as described above is implemented.
[0017] The embodiments of the present application adopting the above technical solutions may include the following advantages: By pre-building a model for supporting cross-modal target semantic retrieval, when content retrieval of a manuscript is required, the target content detection model is used to automatically recall the content medium containing the target semantics. Compared with the manual recognition and retrieval method, the efficiency of content retrieval of manuscript resources is greatly improved, and at the same time, the recall and fallback capabilities for malicious user activities and any newly added target content category samples are improved. Moreover, the target content detection model is not limited to the modality of the manuscript resources. Whether it is video, image, or text modality manuscript content, it can be retrieved and recalled through this method. Description of the Drawings
[0018] The accompanying drawings exemplarily illustrate embodiments and form part of the specification, and are used together with the written description of the specification to explain the exemplary implementation manners of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0019] Figure 1 Schematically shows an operating environment diagram of a manuscript content retrieval method according to Embodiment 1 of the present application; Figure 2 Schematically shows a flowchart of a manuscript content retrieval method according to Embodiment 1 of the present application; Figure 3 Schematically shows a framework structure diagram for training a target content retrieval model; Figure 4 Schematically shows schematic diagrams of various different application scenarios of a target content retrieval model; Figure 5A Schematically shows a schematic diagram of a picture recalled by a target content retrieval model; Figure 5B Schematically shows another schematic diagram of a picture recalled by a target content retrieval model; Figure 5C Schematically shows another schematic diagram of a picture recalled by a target content retrieval model; Figure 6 Schematically shows a block diagram of a manuscript content retrieval device according to Embodiment 2 of the present application; and Figure 7 Schematically shows a schematic diagram of the hardware architecture of a computer device according to Embodiment 3 of the present application. Detailed implementation manners
[0020] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0021] It should be noted that in the embodiments of the present application, the descriptions involving "first", "second", etc. are only for descriptive purposes, and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0022] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and to distinguish each step. Therefore, it should not be construed as a limitation to the present application.
[0023] First, the following provides the term explanations involved in the present application: Query: Query set, and the user input can be videos, images, texts, etc.
[0024] Gallery: The constructed target manuscript index pool, such as the violation sample set, which contains various styles such as videos, images, and texts.
[0025] Open cross-modal retrieval: Based on the Query, query the target semantic content in the Gallery.
[0026] Recall ability: In a search engine, the recall ability is used to measure the ability of the search engine to return relevant documents or web pages under a given query condition. A high recall rate can avoid missing important information and thus improve the user experience.
[0027] Vertical model: Refers to a machine learning model designed for a specific vertical field or industry. These models use deep learning and big data analysis technologies to highly specialize in the processing and prediction of data in a specific field, so as to be able to more accurately solve complex problems in that field.
[0028] Cross-modal retrieval: Refers to establishing a corresponding relationship between two or more modalities (such as text, image, audio, video, etc.), enabling users to retrieve information of another modality through the information of one modality.
[0029] Secondly, to facilitate those skilled in the art to understand the technical solutions provided by the embodiments of the present application, the following explains the related technologies: In the prior art, generally, a unified image-text encoding model is trained based on video / image-text pair data to obtain a trained image-text encoding model to support functions such as text-to-image, image-to-text, text-to-text, and image-to-image. The image-text encoding model obtained through the above solution can only support a relatively coarse-grained semantic understanding of the input data, and many detailed information cannot be covered. For example, some content information hidden in a nine-square grid image, and when the input information for query contains too many details, it will also be ignored.
[0030] To this end, an embodiment of the present application provides a technical solution for retrieving manuscript content. In this technical solution, a manuscript to be processed and a target content detection model are obtained; wherein, the manuscript to be processed is any one of text, image, and video, and the target content detection model is used to support cross-modal target semantic retrieval; through the target content detection model, semantic analysis is performed on the manuscript to be processed from multiple preset dimensions to obtain feature information of multiple dimensions; compliance detection is performed according to the feature information of multiple dimensions to obtain a prediction probability, and a manuscript retrieval result is output according to the prediction probability. By pre-constructing a model for supporting cross-modal target semantic retrieval, when content retrieval of a manuscript is required, the target content detection model is used to automatically recall the content medium containing the target semantics. Compared with the manual recognition and retrieval method, the efficiency of content retrieval of manuscript resources is greatly improved, and at the same time, the recall and fallback capabilities for malicious user activities and arbitrarily added target content category samples are improved. Moreover, the target content detection model is not limited to the modality of the manuscript resources. Whether it is a video, image, or text modality manuscript content, it can be retrieved and recalled through this method. See the following for details.
[0031] Finally, for ease of understanding, an exemplary operating environment is provided below.
[0032] As Figure 1 shown, the environmental schematic diagram includes a service platform 2, a network 4, and a client 6, where: The service platform 2 can be composed of a single or multiple computing devices. The multiple computing devices can include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing device can load a virtual machine based on a virtual image and / or other data that defines a specific software (e.g., operating system, dedicated application, server) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices. A hypervisor can be implemented to manage the use of different virtual machines on the same computing device.
[0033] The service platform 2 can be configured to communicate with the client 6 etc. via the network 4. The network 4 includes various network devices such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network 4 can include physical links such as coaxial cable links, twisted pair cable links, fiber optic links and combinations thereof, or wireless links such as cellular links, satellite links, Wi-Fi links, etc.
[0034] The service platform 2 can provide services such as storage, reading, writing, querying, deleting, etc., such as providing a manuscript content retrieval service for the client.
[0035] The client 6 can be an electronic device running an operating system such as Windows, Android™, or iOS, such as a smart phone, a tablet device, a laptop computer, a virtual reality device, a gaming device, a set-top box, a vehicle-mounted terminal, a smart TV. Based on the above operating systems, various application programs can be run, such as manuscript content retrieval.
[0036] The client 6 can provide / configure a user access page for manipulating the service platform 2 or uploading an object, etc.
[0037] It should be noted that the above devices are exemplary, and in different scenarios or according to different requirements, the number and types of devices can be adjusted.
[0038] Hereinafter, taking the service platform 2 as the execution subject, the technical solutions of the present application will be introduced through multiple embodiments. It should be noted that these embodiments can be implemented in a variety of different forms and should not be construed as being limited only to the embodiments described herein.
[0039] Embodiment 1 Figure 2 A flowchart of a manuscript content retrieval method according to Embodiment 1 of the present application is schematically shown.
[0040] As Figure 2 shown, the manuscript content retrieval method may include steps S202 to S206, wherein: Step S202, obtaining a manuscript to be processed and a target content detection model; wherein, the manuscript to be processed is any one of text, image, and video, and the target content detection model is used to support cross-modal target semantic retrieval; Among them, the manuscript to be processed can be a manuscript published by a user on a content creation platform. After receiving the manuscript published by the user, the content creation platform needs to perform content retrieval on these manuscripts. For example, it retrieves whether the manuscript contains non-compliant content. In this embodiment, since the target content detection model is used to support cross-modal target semantic retrieval, the modality of the manuscript to be processed that needs to be input into the target content detection model for detection is not specifically limited, and the manuscript to be processed can be any one of text, image, and video modalities.
[0041] Step S204, through the target content detection model, perform semantic analysis on the manuscript to be processed from multiple preset dimensions to obtain feature information of multiple dimensions; Among them, the preset dimension refers to multiple information dimensions set in advance, which are used to mine the deep detailed information of the data content. As an example, the preset dimension can include a detection head dimension, a classification head dimension, and a vector head dimension.
[0042] In this embodiment, through the target content detection model, semantic analysis is performed on the manuscript to be processed from multiple preset dimensions to obtain feature information of multiple dimensions. Obtaining the feature information of the manuscript to be processed from multiple dimensions enables the obtained feature information to cover more detailed information in the manuscript to be processed, that is, to perform a finer-grained semantic understanding of the manuscript to be processed. Among them, the feature information of multiple dimensions can include the feature information of the detection head dimension, the feature information of the classification head dimension, and the feature information of the vector head dimension. The feature information of the detection head dimension and the feature information of the classification head dimension are used to mine the fine-grained semantic understanding ability, and the feature information of the vector head dimension is used to align the graph-text recall ability.
[0043] Step S206, perform compliance detection according to the feature information of multiple dimensions to obtain a prediction probability, and output a manuscript retrieval result according to the prediction probability.
[0044] In this embodiment, after extracting the feature information of multiple dimensions in the manuscript to be processed, compliance detection can be further performed according to the extracted feature information of multiple dimensions to obtain a prediction probability. Specifically, a Gallery containing a negative sample dataset can be pre-maintained. By extracting the feature information of each negative sample data in the Gallery, then calculating the similarity between the manuscript to be processed and each negative sample data in the Gallery according to the extracted feature information, and outputting the prediction probability. This prediction probability can represent the probability that the target content exists in the manuscript to be processed. Therefore, the content review platform can output the manuscript retrieval result of the manuscript to be processed according to this prediction probability. If the manuscript to be processed fails the content retrieval, the manuscript to be processed is rejected. At this time, other users in the content creation platform cannot see the manuscript to be processed; if the manuscript to be processed passes the content retrieval, the manuscript to be processed is published to the content creation platform. At this time, other users in the content creation platform can see the manuscript to be processed.
[0045] In an alternative embodiment of the present application, the method further includes: When the prediction probability is greater than a preset threshold, output the manuscript retrieval result that the manuscript to be processed fails the retrieval.
[0046] Wherein, the preset threshold is a pre-set maximum critical probability value for judging whether the manuscript to be processed is compliant. When the prediction probability is greater than the preset threshold, it is considered that the manuscript to be processed is non-compliant, and the manuscript retrieval result that the manuscript to be processed fails the retrieval can be output. The manuscript retrieval result describes the non-compliant items for which the manuscript to be processed fails the retrieval. For example, the manuscript retrieval result is "There is a certain violation item in the XXX area of the manuscript", "There is a certain violation item in the XXX area of the manuscript", etc.
[0047] In an alternative embodiment of the present application, the method further includes: When the manuscript retrieval result belongs to a preset type, perform a secondary retrieval on the manuscript to be processed through a preset retrieval channel to obtain a target manuscript retrieval result.
[0048] In this embodiment, for some manuscripts containing high-risk target content, they can be directly rejected after failing the retrieval by the target content detection model; while for some manuscripts containing low-risk target content, after failing the retrieval by the target content detection model, a secondary retrieval can also be performed on the manuscript to be processed through a preset retrieval channel to obtain a target manuscript retrieval result. When the target manuscript retrieval result is non-compliant, the manuscript to be processed is rejected, and when the target manuscript retrieval result is compliant, the manuscript to be processed is published to the content creation platform.
[0049] Specifically, by determining whether the manuscript retrieval result belongs to a preset type, which may be a low-risk violation type. In the case where the manuscript retrieval result belongs to the preset type, a secondary retrieval is performed on the manuscript to be processed through a preset retrieval channel to obtain a target manuscript retrieval result. The preset retrieval channel may be a preset manual retrieval channel, or the preset retrieval channel may also be a coarse-grained content review model. The embodiments of the present application do not make specific limitations in this regard.
[0050] In an alternative embodiment of the present application, before step 202, the method may further include the following steps: Obtain positive sample data and negative sample data in multiple modalities; perform feature extraction on the positive sample data in multiple modalities to obtain first feature information, and perform feature extraction on the negative sample data in multiple modalities to obtain second feature information; wherein, the first feature information and the second feature information belong to a first granularity; train an initial content detection model according to the first feature information and the second feature information.
[0051] In this embodiment, positive sample data and negative sample data in multiple modalities are obtained for training the initial content detection model, and the initial content detection model belongs to a coarse-grained detection model. Among them, the modalities of the positive sample data and the negative sample data can be any one of text, image, and video.
[0052] After obtaining positive sample data and negative sample data in multiple modalities, perform feature extraction on the positive sample data in multiple modalities to obtain first feature information, and perform feature extraction on the negative sample data in multiple modalities to obtain second feature information; wherein, the first feature information and the second feature information belong to a first granularity, and this first granularity is the coarse granularity; finally, train an initial content detection model according to the first feature information and the second feature information.
[0053] In an alternative embodiment of the present application, the method further includes: Perform semantic analysis on the positive sample data in multiple modalities to obtain third feature information in multiple dimensions, and perform semantic analysis on the negative sample data in multiple modalities to obtain fourth feature information in multiple dimensions; wherein, the third feature information and the fourth feature information belong to a second granularity; transform the third feature information and the fourth feature information into a preset vector space; train the initial content detection model according to the transformed third feature information and fourth feature information to obtain the target content detection model.
[0054] In this embodiment, the positive sample data can be used as the input query set Query, and the negative sample data can be used as the violation sample set Gallery. Through semantic analysis of the positive sample data in multiple modalities, third feature information in multiple dimensions is obtained, and through semantic analysis of the negative sample data in multiple modalities, fourth feature information in multiple dimensions is obtained; wherein, the third feature information and the fourth feature information belong to the second granularity, and this second granularity is the fine granularity.
[0055] After the third feature information and the fourth feature information are extracted, by transforming the third feature information and the fourth feature information into a preset vector space, the third feature information and the fourth feature information are normalized to achieve the purpose of uniformly encoding the data in multiple modalities. Finally, according to the transformed third feature information and fourth feature information, the initial content detection model is trained until the prediction effect of the initial content detection model reaches the expectation, and then the target content detection model is output.
[0056] In the embodiment, the third feature information and the fourth feature information respectively include: feature information in the detection head dimension, feature information in the classification head dimension, and feature information in the vector head dimension; wherein, the feature information in the detection head dimension and the feature information in the classification head dimension are used to mine the semantic understanding ability of the second granularity, and the feature information in the vector head dimension is used to align the image-text recall ability.
[0057] In an optional embodiment of the present application, the training of the initial content detection model according to the transformed third feature information and fourth feature information to obtain the target content detection model includes: Determine the similarity between the transformed third feature information and the fourth feature information; determine multiple model losses according to the similarity according to multiple preset loss functions; optimize the parameters of the initial content detection model according to the multiple model losses to obtain the target content detection model.
[0058] In this embodiment, by calculating the similarity between the transformed third feature information and the fourth feature information, and then determining multiple model losses according to the similarity according to multiple preset loss functions, that is, a model loss is calculated for the feature information in the detection head dimension, a model loss is calculated for the feature information in the classification head dimension, and a model loss is calculated for the feature information in the vector head dimension. Finally, the parameters of the initial content detection model are optimized according to these model losses until the prediction effect of the initial content detection model reaches the expectation, and then the target content detection model is output. Figure 3It shows a framework structure diagram for training a target content retrieval model. Specifically, there are two input modalities in the training process: Captions (text descriptions) and Images. Among them, for Captions, text-encoder text information can be directly extracted for training. For Images, head information is extracted (such as embedding-head (vector head), detection / classification head information, norm-view standard view information, image-encoder image encoding information), and finally the extracted information is transmitted to the contrastive learning module for training to obtain a target content detection model.
[0059] In this embodiment, the target content detection model includes the ability to mine and understand details and the ability to understand the overall content; it supports the recognition / detection of small targets in the manuscript. At the same time, a unified encoder is used on the text and image sides, supporting cross-modal, which makes the model lightweight. Moreover, the model has strong scalability and can be used for further adjustment (finetune) of downstream services. For example, a vertical model can also be trained based on the target content detection model, including: text vertical model, multi-modal vertical model, image vertical model, large language model, etc. Figure 4 It shows schematic diagrams of various different application scenarios of a target content retrieval model. Among them, for a large language model, by inputting a question, the model can understand this question and give an answer. For example, when the question is "Describe this picture", after the large language model recognizes the content in the picture, it outputs the content detected from the picture. Figures 5A - 5C Schematic diagram of the picture recalled by the target content retrieval model, where Figure 5A It is a recall for a specific comic, Figure 5B It is a recall based on the text description "Steps of cooking", Figure 5C It is a recall based on the text description "Scene of playing on the beach".
[0060] In this embodiment, by training a target content detection model that supports cross-modal target semantic retrieval, the target content detection model is used to automatically recall content media containing the target semantics. Compared with the method of manual recognition and retrieval, the efficiency of manuscript content retrieval is greatly improved. At the same time, the recall and fallback ability for malicious user activities and any newly added target content category samples is enhanced. The model is not limited to the modality of the input manuscript, and can retrieve and recall the content of manuscripts in video, image, or text modality through this method.
[0061] Embodiment 2 Figure 3Schematically shown is a block diagram of a manuscript content retrieval device according to Embodiment 2 of the present application. The device can be divided into one or more program modules. One or more program modules are stored in a storage medium and executed by one or more processors to complete the embodiments of the present application. The program modules referred to in the embodiments of the present application refer to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. As Figure 3 shown, the manuscript content retrieval device 300 may include: a manuscript to be processed acquisition module 301, a semantic analysis module 302, and a compliance detection module 303, where: The manuscript to be processed acquisition module 301 is used to acquire a manuscript to be processed and a target content detection model; wherein, the manuscript to be processed is any one of text, image, and video, and the target content detection model is used to support cross-modal target semantic retrieval; The semantic analysis module 302 is used to perform semantic analysis on the manuscript to be processed from multiple preset dimensions through the target content detection model to obtain feature information of multiple dimensions; The compliance detection module 303 is used to perform compliance detection based on the feature information of multiple dimensions to obtain a prediction probability, and output a manuscript retrieval result according to the prediction probability.
[0062] In an alternative embodiment of the present application, the device further includes: A retrieval result output module, which is used to output that the manuscript retrieval result is that the manuscript to be processed fails the retrieval when the prediction probability is greater than a preset threshold.
[0063] In an alternative embodiment of the present application, the device further includes: A secondary retrieval module, which is used to perform secondary retrieval on the manuscript to be processed through a preset retrieval channel to obtain a target manuscript retrieval result when the manuscript retrieval result belongs to a preset type.
[0064] In an alternative embodiment of the present application, the device further includes: A positive sample data acquisition module, which is used to acquire positive sample data and negative sample data of multiple modalities; A feature extraction module, which is used to extract first feature information from the positive sample data of multiple modalities and extract second feature information from the negative sample data of multiple modalities; wherein, the first feature information and the second feature information belong to the first granularity; An initial content detection model training module, which is used to train an initial content detection model according to the first feature information and the second feature information.
[0065] In an alternative embodiment of the present application, the device further includes: A semantic analysis module, configured to perform semantic analysis on the positive sample data of the multiple modalities to obtain third feature information of multiple dimensions, and perform semantic analysis on the negative sample data of the multiple modalities to obtain fourth feature information of multiple dimensions; wherein, the third feature information and the fourth feature information belong to the second granularity; A feature information beautification module, configured to transform the third feature information and the fourth feature information into a preset vector space; A target content detection model training module, configured to train the initial content detection model according to the transformed third feature information and fourth feature information to obtain the target content detection model.
[0066] In an alternative embodiment of the present application, the target content detection model training module includes: A similarity determination sub-module, configured to determine the similarity between the transformed third feature information and fourth feature information; A model loss determination sub-module, configured to determine multiple model losses according to the similarity according to multiple preset loss functions; A parameter optimization module of the model, configured to optimize the parameters of the initial content detection model according to the multiple model losses to obtain the target content detection model.
[0067] In an alternative embodiment of the present application, the third feature information and the fourth feature information respectively include: feature information of the detection head dimension, feature information of the classification head dimension, and feature information of the vector head dimension; the feature information of the detection head dimension and the feature information of the classification head dimension are used to mine the semantic understanding ability of the second granularity, and the feature information of the vector head dimension is used to align the image-text recall ability.
[0068] Embodiment III Figure 4 FIG. schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing a manuscript content retrieval method according to Embodiment III of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. As Figure 4 shown, the computer device 10000 includes but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the manuscript content retrieval method. In addition, the memory 10010 may also be used to temporarily store various data that have been output or will be output.
[0069] In some embodiments, the processor 10020 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other chip. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.
[0070] The network interface 10030 may include a wireless network interface or a wired network interface, which is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an enterprise intranet (Intranet), the Internet, Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi, etc.
[0071] It should be noted that Figure 4 only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be alternatively implemented.
[0072] In this embodiment, the manuscript content retrieval method stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.
[0073] Embodiment 4 The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the manuscript content retrieval method in the embodiment are implemented.
[0074] In this embodiment, the computer-readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (e.g., SD or DX memories, etc.), random access memories (RAM), static random access memories (SRAM), read-only memories (ROM), electrically erasable programmable read-only memories (EEPROM), programmable read-only memories (PROM), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the manuscript content retrieval method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or will be output.
[0075] Embodiment 5 The embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor implements the method in the above embodiment.
[0076] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device. Thus, they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0077] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A method for retrieving manuscript content, characterized in that, The method comprises: Obtaining a manuscript to be processed and a target content detection model; wherein the manuscript to be processed is in any modality of text, image, and video, and the target content detection model is used to support cross-modal target semantic retrieval; By using the target content detection model, semantic analysis is performed on the manuscript to be processed from multiple preset dimensions to obtain feature information of multiple dimensions; Compliance detection is performed based on the feature information of the multiple dimensions to obtain a predicted probability, and a manuscript retrieval result is output based on the predicted probability.
2. The manuscript content retrieval method according to claim 1, characterized in that Outputting the manuscript search results according to the predicted probability includes: When the predicted probability is greater than a preset threshold, the manuscript search result is output as the manuscript to be processed failing the search.
3. The manuscript content retrieval method according to claim 2, characterized in that, The method further comprises: In the case that the manuscript retrieval result belongs to a preset type, a secondary retrieval is performed on the manuscript to be processed through a preset retrieval channel to obtain a target manuscript retrieval result.
4. The manuscript content retrieval method according to claim 1, characterized in that, Before the step of obtaining the manuscript to be processed and the target content detection model, the method further includes: Obtain positive and negative sample data of multiple modalities; Performing feature extraction on the positive sample data of the multiple modalities to obtain first feature information, and performing feature extraction on the negative sample data of the multiple modalities to obtain second feature information; wherein the first feature information and the second feature information belong to a first granularity; An initial content detection model is trained based on the first feature information and the second feature information.
5. The manuscript content retrieval method according to claim 4, characterized in that The method further comprises: Semantically analyzing the positive sample data of the multiple modalities to obtain third feature information of multiple dimensions, and semantically analyzing the negative sample data of the multiple modalities to obtain fourth feature information of multiple dimensions; wherein the third feature information and the fourth feature information belong to the second granularity; Converting the third feature information and the fourth feature information into a preset vector space; The initial content detection model is trained based on the converted third feature information and fourth feature information to obtain the target content detection model.
6. The manuscript content retrieval method according to claim 5, characterized in that The training of the initial content detection model according to the converted third feature information and fourth feature information to obtain the target content detection model includes: Determining the similarity between the converted third feature information and the fourth feature information; Determining multiple model losses based on the similarities according to multiple preset loss functions; Parameters of the initial content detection model are optimized according to the multiple model losses to obtain the target content detection model.
7. The manuscript content retrieval method according to claim 5 or 6, characterized in that The third feature information and the fourth feature information respectively include: feature information of the detection head dimension, feature information of the classification head dimension and feature information of the vector head dimension; the feature information of the detection head dimension and the feature information of the classification head dimension are used to mine the semantic understanding ability of the second granularity, and the feature information of the vector head dimension is used to align the image and text recall ability.
8. A manuscript content retrieval device, characterized in that, The device comprises: A module for acquiring manuscripts to be processed, configured to acquire manuscripts to be processed and a target content detection model; wherein the manuscripts to be processed are in any of the modalities of text, image, and video, and the target content detection model is configured to support cross-modal target semantic retrieval; A semantic analysis module, configured to perform semantic analysis on the to-be-processed manuscript from multiple preset dimensions through the target content detection model to obtain feature information of multiple dimensions; A compliance detection module, configured to perform compliance detection based on the feature information of multiple dimensions to obtain a prediction probability, and output a manuscript retrieval result according to the prediction probability.
9. A computer device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein: The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the manuscript content retrieval method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the manuscript content retrieval method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the manuscript content retrieval method according to claims 1 to 7 are implemented.
Citation Information
Patent Citations
Model training method and device, computer equipment and storage medium
CN117033720A
Resource auditing method and device
CN118400573A
Instance level scene recognition with a vision language model
US11978271B1
Method and system of using domain specific knowledge in retrieving multimodal assets
US20240248901A1
Cross-modal retrieval system and method based on pre-training model and recall and ranking
WO2023065617A1