Cross-modal retrieval method, device, medium and equipment for MOOCs graphic and text resources
By building a fusion model of image and text content features and search intent features on the MOOCs platform and using hash coding matrix for indexing, the problems of inaccurate retrieval results and high computational complexity in existing technologies are solved, and efficient and accurate image and text resource retrieval is achieved.
Patent Information
- Application Number
- CN202510936528.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing cross-modal retrieval technology for images and texts on MOOCs platforms cannot effectively capture learners' search intentions, resulting in inaccurate retrieval results and high computational complexity.
By building a cross-modal retrieval model that integrates the content features of images and text with search intent features, and using hash coding matrix for indexing, efficient matching of image and text resources can be achieved.
It improves the accuracy and efficiency of retrieval, reduces computational complexity, meets the complex and changing learning needs of the MOOCs platform, and improves user experience.
Smart Images

Figure CN120429453B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information retrieval technology, and in particular to a method, apparatus, medium, and device for cross-modal retrieval of MOOCs graphic and text resources. Background Art
[0002] In the digital education era, Massive Open Online Courses (MOOCs) learning platforms have attracted a large number of users with their in-depth teaching content and accessible learning methods. Learners can learn anytime and anywhere, breaking the constraints of time and space. Numerous MOOC platforms have emerged both domestically and internationally, such as China University MOOC, XuetangX, Udemy, Coursera, and edX, offering a wealth of course content and learning resources. MOOC platforms typically include a variety of learning resources, including videos, handouts, images, text, and audio. Images and text, as core knowledge carriers, play a crucial role. To help learners access the content they need more efficiently, a large number of cross-modal image-text retrieval requirements have emerged on MOOC platforms, such as searching for images using text or searching for related text using images. This demand has driven the development of cross-modal retrieval technology.
[0003] Among related technologies, existing cross-modal image-text retrieval technology mainly relies on semantic feature matching of image and text resources. Although it can complete basic retrieval tasks to a certain extent, it still faces many major problems. On the one hand, image and text resources belong to the visual modality and language modality respectively. There is a semantic gap between the two, which will lead to the inability to fully align the image and text, affecting the retrieval effect. On the other hand, existing methods ignore the learners' search intentions and cannot effectively capture the real needs behind them. When learners search for resources on MOOCs platforms, they often express different intentions based on their learning goals, cognitive needs, and interaction methods. The existing retrieval methods based on semantic matching are difficult to meet these personalized needs, resulting in inaccurate retrieval results. Summary of the Invention
[0004] The embodiments of the present disclosure at least provide a method, apparatus, medium, and device for cross-modal retrieval of MOOCs graphic and text resources. By integrating content features of different modalities with search intent features to construct a retrieval model, hash coding is used to achieve efficient cross-modal retrieval, effectively improving the accuracy and efficiency of retrieval and reducing computational complexity.
[0005] The present disclosure provides a method for cross-modal retrieval of MOOCs graphic and text resources, including:
[0006] Obtain image resource collections and text resource collections from the MOOCs platform;
[0007] Constructing an image content feature matrix and an image search intent feature matrix based on the image resource set, and constructing a text content feature matrix and a text search intent feature matrix based on the text resource set;
[0008] Based on the image content feature matrix, the image search intent feature matrix, the text content feature matrix, and the text search intent feature matrix, a MOOCs resource image-text cross-modal retrieval model is constructed;
[0009] Determine a hash coding matrix based on the MOOCs resource image-text cross-modal retrieval model, construct an image hash function based on the image content feature matrix and the hash coding matrix, and construct a text hash function based on the text content feature matrix and the hash coding matrix;
[0010] The information to be queried is obtained, and a target query result is determined based on the information to be queried, the hash coding matrix, the image hash function, and the text hash function.
[0011] The present disclosure provides a cross-modal retrieval device for MOOCs graphic and text resources, including:
[0012] Resource acquisition module, used to obtain image resource collection and text resource collection of MOOCs platform;
[0013] A matrix construction module, configured to construct an image content feature matrix and an image search intent feature matrix based on the image resource set, and to construct a text content feature matrix and a text search intent feature matrix based on the text resource set;
[0014] A model construction module is used to construct a MOOCs resource image-text cross-modal retrieval model based on the image content feature matrix, the image search intent feature matrix, the text content feature matrix, and the text search intent feature matrix;
[0015] A function construction module is used to determine a hash coding matrix based on the MOOCs resource image-text cross-modal retrieval model, construct an image hash function based on the image content feature matrix and the hash coding matrix, and construct a text hash function based on the text content feature matrix and the hash coding matrix;
[0016] The information query module is used to obtain the information to be queried and determine the target query result based on the information to be queried, the hash coding matrix, the image hash function and the text hash function.
[0017] An embodiment of the present disclosure provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor and the memory communicate via the bus. When the machine-readable instructions are executed by the processor, the cross-modal retrieval method for MOOCs graphic and text resources as described in any of the possible implementation methods described above is performed.
[0018] An embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for cross-modal retrieval of MOOCs graphic and text resources as described in any of the above possible implementations is implemented.
[0019] The MOOCs graphic and text resource cross-modal retrieval method, apparatus, medium, and device provided in the embodiments of the present disclosure improve the accuracy of cross-modal retrieval by introducing the content features of images and texts and combining them with search intent features, and can effectively capture the learners' real search needs. In addition, the use of hash coding matrices for indexing image and text resources not only speeds up the retrieval process but also improves computational efficiency. In this way, more accurate graphic and text resource retrieval can be achieved, effectively improving the accuracy and efficiency of retrieval, reducing computational complexity, meeting the complex and changing learning needs on the MOOCs platform, and improving user experience.
[0020] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings that need to be cited in the embodiments. The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without inventive effort.
[0022] Figure 1 A flowchart of a cross-modal retrieval method for MOOCs graphic and text resources provided by an embodiment of the present disclosure is shown;
[0023] Figure 2 A flowchart of a method for constructing an image hash function and a text hash function provided by an embodiment of the present disclosure is shown;
[0024] Figure 3A flow chart of a method for determining a target query result provided by an embodiment of the present disclosure is shown;
[0025] Figure 4 A schematic diagram of the structure of a cross-modal retrieval device for MOOCs graphic and text resources provided by an embodiment of the present disclosure is shown;
[0026] Figure 5 A schematic structural diagram of a computer device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. The components of the embodiments of the present disclosure generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the disclosure for which protection is sought, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.
[0028] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0029] The term "and / or" herein simply describes an association relationship, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, and the existence of B alone. In addition, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.
[0030] To facilitate understanding of this embodiment, the execution subject of the cross-modal retrieval method for MOOCs graphic and text resources provided by the embodiment of the present disclosure is first introduced in detail. The execution subject of the cross-modal retrieval method for MOOCs graphic and text resources provided by the embodiment of the present disclosure is a computer device. The computer device can be a terminal device or a server. Among them, the terminal device can also be a mobile device, a user terminal, a terminal, a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data and artificial intelligence platforms. Optionally, the method can also be applied to an implementation environment composed of computer devices and servers.
[0031] The following describes in detail the cross-modal retrieval method for MOOCs graphic resources provided by the embodiment of the present application in conjunction with the accompanying drawings. Figure 1 FIG. 1 is a flowchart of a method for cross-modal retrieval of MOOCs graphic and text resources provided by an embodiment of the present disclosure. The method includes the following steps S101 to S105:
[0032] S101, obtaining an image resource collection and a text resource collection of a MOOCs platform.
[0033] As you can understand, MOOCs (Massive Open Online Courses) platforms are virtual platforms that offer online educational courses, providing free or paid course content to learners worldwide. For example, Coursera, edX, and Udemy are all MOOC platforms. These platforms typically offer courses in various disciplines, accompanied by learning resources such as video lectures, textbooks, and quizzes. The image resource collection includes all visual elements related to the platform's courses, such as images, illustrations, screenshots, and video frames. Examples include charts in courses, screenshots from instructional videos, illustrations in courseware, and various visual aids used in course presentations. Image resources refer to all information that can be accessed through visual perception and are typically used to convey concepts, data, or scenarios in the course, helping learners better understand complex content. The text resource collection includes textual materials such as course descriptions, handouts, textbooks, and discussion forum content. These text resources typically include detailed course outlines, teaching content, study guides, and transcripts of class discussions, providing core course knowledge and learning support information. Text resources here refer to all information presented in text form, which can help students understand the theory and background of the course more deeply.
[0034] S102, constructing an image content feature matrix and an image search intent feature matrix based on the image resource set, and constructing a text content feature matrix and a text search intent feature matrix based on the text resource set.
[0035] Specifically, for image resource collections, two types of features need to be extracted: content features and search intent features. Content features describe the visual information of the image itself, such as color, texture, shape, etc., and are usually extracted through deep learning models such as convolutional neural networks. Search intent features attempt to capture the meaning of the image in the search context, such as the relevance of the image to the course topic, the position of the image in the course, etc. These features may need to be extracted in combination with course structure and user behavior data. Furthermore, the extracted features are organized into matrix form, namely the image content feature matrix and the image search intent feature matrix.
[0036] For example, assuming that there are Image resources , the image content encoder is , build image resources Image content feature matrix with dimensional semantic content Similarly, for Image resources , the image search intent encoder is , build image resources The image search intent feature matrix of the dimensional search intent is .
[0037] Similarly, for text resource collections, we also need to extract content features and search intent features. Content features may include word frequency, word vectors, and topic models, describing the semantic information of the text. Search intent features may include the text's position within the course and its relevance to other texts. Similarly, these features are organized into a text content feature matrix and a text search intent feature matrix.
[0038] For example, suppose there is Text resources , the text content encoder is , which can construct text resources Text content feature matrix of dimensional semantic content Similarly, for Text resources , the text search intent encoder is , building a text resource The text search intent feature matrix of the dimension search intent is .
[0039] S103: Constructing a MOOCs resource image-text cross-modal retrieval model based on the image content feature matrix, the image search intent feature matrix, the text content feature matrix, and the text search intent feature matrix.
[0040] It's understandable that after obtaining the image content feature matrix, image search intent feature matrix, text content feature matrix, and text search intent feature matrix, we can use these four feature matrices to build a cross-modal retrieval model. The goal of this model is to learn the mapping relationship between images and text, so that when given a query (whether image or text), it can find the most relevant resources in the other modality.
[0041] Here, the MOOCs resource image and text cross-modal retrieval model can be expressed as:
[0042] ;
[0043] in, Expressed as a cross-modal retrieval objective function; is expressed as a weight coefficient, and , (i=1,2,…,7), ; Represented as image content feature matrix; Represented as the basis vector of the image content feature matrix; Represents the potential shared content features of image resources and text resources; F is represented as the square root of the sum of the squares of all elements of the matrix; Represented as a text content feature matrix; Represented as the basis vector of the text content feature matrix; Represented as a potentially shared semantic description The projection matrix on the hash space; B is represented as the hash encoding matrix; Represented as an image search intent feature matrix; Represented as the basis vector of the image search intent feature matrix; Representing the potential shared search intent features of image resources and text resources; Represented as a text search intent feature matrix; Represented as the basis vector of the text search intent feature matrix; Represented as potential shared search intent features Projection matrix onto hash space.
[0044] S104: determining a hash coding matrix based on the MOOCs resource image-text cross-modal retrieval model, constructing an image hash function based on the image content feature matrix and the hash coding matrix, and constructing a text hash function based on the text content feature matrix and the hash coding matrix.
[0045] Specifically, hash coding is a technique for mapping high-dimensional data into a low-dimensional space in order to accelerate retrieval and improve computational efficiency. The main purpose of hash coding is to reduce data storage and computational costs by converting the original image and text features into low-dimensional binary hash codes, while quickly matching similar images and texts during the retrieval process. For the cross-modal retrieval task of images and texts in MOOCs resources, in order to compress and map the above-mentioned image content feature matrix and text content feature matrix into the same hash space, a hash coding matrix is required. It can learn effective image and text feature mappings through an optimization process. Here, the constructed cross-modal retrieval model will be used to determine a hash coding matrix, which will be used to map image and text features into the hash space.
[0046] Here, the hash coding matrix of the MOOCs resource image and text cross-modal retrieval model can be learned through an iterative optimization method, which can specifically include the following steps (1) to (4):
[0047] (1) Random initialization 、 、 、 、 、 、 、 , B;
[0048] (2) Fix the other variables in turn and update the single variable according to the following rules:
[0049] The update expression is expressed as:
[0050] ;
[0051] Where I represents the identity matrix; Expressed as The transposed matrix of
[0052] The update expression is expressed as:
[0053] ;
[0054] The update expression is expressed as:
[0055] ;
[0056] The update expression is expressed as:
[0057] ;
[0058] The update expression is expressed as:
[0059] ;
[0060] The update expression is expressed as:
[0061] ;
[0062] The update expression is expressed as:
[0063] ;
[0064] The update expression is expressed as:
[0065] ;
[0066] The update expression of B is expressed as:
[0067] ;
[0068] (3) Based on the updated 、 、 、 、 、 、 、 , B. solving the MOOCs resource graphic and text cross-modal retrieval model;
[0069] (4) Repeat steps (2) to (3) until the number of iterations reaches a preset value and / or the cross-modal retrieval objective function converges, and determine the hash coding matrix.
[0070] Here, after initializing the relevant parameters in the hash coding matrix (such as the weight matrix and feature matrix), some variables are fixed in sequence, and other variables are updated according to specific update rules. The specific update expression includes iterative updates to various variables in the hash coding matrix, such as the basis vectors of the image content feature matrix / text content feature matrix, the potential shared content features / potential shared search intent features of image and text resources, and the hash matrix. Each iteration gradually optimizes these matrices, and based on these updated variables, a cross-modal retrieval model for MOOC resources, including images and text, can be solved.
[0071] Here, the solution process is essentially to use the information carried by the updated variables to further optimize the overall performance of the model. and By mapping image content features from different perspectives, they can more accurately convert image features into hash space; and It captures the potential shared content features and potential shared search intent features between images and text respectively, which enables the model to better understand the semantic relationship between the two modalities; and Mapping the search intent features of images and texts enhances the model's understanding of user search intent, making images and texts with similar search intent closer in the hash space. and As a projection matrix, α projects shared features into the hash space, ensuring that key semantic information is preserved during the hash encoding process. B, the hash encoding matrix, is the key to ultimately mapping high-dimensional features into the low-dimensional hash space. By substituting these updated variables into the model, the model can readjust the position and relationship of images and text in the hash space based on the new parameter information, bringing semantically similar images and text closer together in the hash space, thereby improving the accuracy of cross-modal retrieval.
[0072] Furthermore, steps (3) to (4) are repeated until the objective function converges or the set number of iterations is reached, and the hash coding matrix is obtained. In each iteration, some variables are fixed in turn, and then other variables are updated according to the update rules in step (2). As the iteration proceeds, each round will gradually optimize the variables in the hash coding matrix, and the objective function of the model is also constantly adjusted. The objective function contains information such as the distance measurement or correlation measurement between the image and the text. By minimizing the objective function, the image and text can be made as close as possible to their corresponding similar content in the hash space. When the number of iterations reaches the preset value, it means that the model has undergone enough optimization steps; or when the cross-modal retrieval objective function converges, it means that the performance of the model has stabilized, and further iterations will have limited improvement on the model performance. At this point, the final hash coding matrix can be determined. This matrix can effectively map the image and text resources in the MOOCs platform to a low-dimensional hash space, achieving efficient cross-modal retrieval.
[0073] The hash coding matrix finally determined in the embodiment of the present disclosure, because it incorporates the information of the search intent feature matrix, can more effectively map the image and text resources in the MOOCs platform to a low-dimensional hash space. When facing the actual search needs of users, it can provide more accurate and user-expected cross-modal retrieval results, greatly improving the practicality of the retrieval and user satisfaction.
[0074] It can be understood that after obtaining the hash coding matrix, an image hash function can be constructed based on the image content feature matrix and the hash coding matrix, and a text hash function can be constructed based on the text content feature matrix and the hash coding matrix. Constructing an image hash function based on the image content feature matrix and the hash coding matrix essentially involves using the hash coding matrix to perform a linear or nonlinear transformation on the image content feature matrix. Through this transformation, the high-dimensional feature vectors in the image content feature matrix are converted into low-dimensional hash codes, significantly reducing the feature dimensionality while retaining the key semantic information of the image, thereby improving the efficiency of subsequent retrieval. For example, in a MOOC course, an image containing a complex circuit diagram has an image content feature matrix that records information such as the direction of the circuit lines, the shape and color of the components. By operating with the hash coding matrix, these high-dimensional features are converted into concise hash codes, facilitating rapid matching in the hash space. Similarly, the hash coding matrix is used to transform the text content feature matrix, mapping the high-dimensional text feature vectors into a low-dimensional hash space to generate a hash code for the text. Taking the description text of a MOOC course as an example, the text content feature matrix records information such as the course topic and the knowledge points involved. After operation with the hash coding matrix, it is converted into a hash code, so that the similarity can be quickly compared with the hash codes of other related texts or images during retrieval.
[0075] For example, referring to Figure 2 As shown, in the process of constructing the image hash function and the text hash function, the following steps S201 to S203 may be included:
[0076] S201 : constructing an initial image hash function based on the image content feature matrix and the hash coding matrix; and constructing an initial text hash function based on the text content feature matrix and the hash coding matrix.
[0077] Here, the image content feature matrix stores the feature information of image resources on the MOOCs platform. These features are high-dimensional vectors extracted from images using specific image feature extraction methods. The hash code matrix, previously learned through iterative optimization methods, is a key matrix used to map high-dimensional features to a low-dimensional hash space. By performing specific mathematical operations, such as matrix multiplication, on the image content feature matrix and the hash code matrix, an initial image hash function can be constructed. This initial function provides a preliminary rule for mapping image features to the hash space. Similarly, the text content feature matrix contains feature vectors obtained after feature extraction of text such as course descriptions and handouts. By performing corresponding operations on the text content feature matrix and the hash code matrix, an initial text hash function can be constructed, laying the foundation for mapping text features to the hash space.
[0078] S202, constructing a hash target function based on the initial image hash function and the initial text hash function; and solving an image parameter matrix corresponding to the initial image hash function and a text parameter matrix corresponding to the initial text hash function based on the hash target function.
[0079] Specifically, the hash objective function is constructed to measure the performance of the initial image hash function and the initial text hash function in cross-modal retrieval tasks. This objective function usually takes into account factors such as the similarity and distance measurement of images and texts in the hash space. For example, it is hoped that the hash codes of semantically similar images and texts in the hash space have a smaller Hamming distance, while the hash codes of semantically dissimilar images and texts have a larger Hamming distance. By optimizing the hash objective function, the image parameter matrix corresponding to the initial image hash function and the text parameter matrix corresponding to the initial text hash function can be solved respectively. The role of these parameter matrices is to adjust and optimize the initial hash function so that the hash function can more accurately map image and text features to the hash space to meet the needs of cross-modal retrieval.
[0080] Here, the initial image hash function can be expressed as:
[0081] ;
[0082] The initial text hash function can be expressed as:
[0083] ;
[0084] The hash objective function can be expressed as:
[0085] ;
[0086] in, It is represented as a hash objective function; n represents the number of resources; Represented as an image parameter matrix; Represented as a text parameter matrix; Represented as image content feature matrix The feature vector corresponding to the i-th image resource; Represented as the i-th column of the hash coding matrix; Represented as a text content feature matrix The feature vector corresponding to the i-th text resource.
[0087] Exemplarily, when respectively solving the image parameter matrix corresponding to the initial image hash function and the text parameter matrix corresponding to the initial text hash function based on the hash objective function, the following (a) to (d) may be included:
[0088] (a) Initialization and ;
[0089] (b) Update separately by gradient descent method and ,include:
[0090] ;
[0091] ;
[0092] in, Expressed as element-wise multiplication;
[0093] (c) Based on the updated and Solving the hash objective function;
[0094] (d) Repeat steps b to c until the hash objective function converges and determine and .
[0095] Specifically, in and After initialization, the gradient descent method is used to update and , The step size of the gradient descent determines the magnitude of each update. If the step size is too large, the model may not converge; if the step size is too small, the convergence speed will be slow. The core idea of the gradient descent method is to update the parameters along the negative gradient direction of the objective function so that the value of the objective function gradually decreases. Then, based on the updated and Solve the hash objective function. Substitute the updated parameters into the hash objective function and calculate the objective function value under the current parameters to evaluate the effect of the parameter update. Finally, repeat steps (b) to (c) until the hash objective function converges. When the value of the objective function changes very little or reaches the preset number of iterations, the model is considered to have converged, and the and The two parameter matrices will be used to construct image hash functions and text hash functions, thereby achieving an effective mapping of image and text features to hash space in MOOCs resource cross-modal retrieval.
[0096] S203 : Determine the image hash function based on the image content feature matrix, the image parameter matrix, and the initial image hash function; and determine the text hash function based on the text content feature matrix, the text parameter matrix, and the initial text hash function.
[0097] It is understandable that after obtaining the image parameter matrix, it can be combined with the image content feature matrix and the initial image hash function. In this way, the final image hash function can more accurately convert image features into hash codes. Similarly, for the text hash function, the text parameter matrix is correspondingly calculated and combined with the text content feature matrix and the initial text hash function, and the initial text hash function is optimized to determine the final text hash function to achieve accurate conversion of text features to hash codes. Through these two final hash functions of different modalities, the image and text resources in the MOOCs platform can be mapped to the hash space respectively, providing efficient data representation for subsequent cross-modal retrieval.
[0098] S105 , obtaining information to be queried, and determining a target query result based on the information to be queried, the hash coding matrix, the image hash function, and the text hash function.
[0099] Here, the query information can be an image uploaded by the user or text entered, and can include either image or text query information. For image queries, an image hash function is used to convert the query information into a hash code; for text queries, a text hash function is used. The Hamming distance or cosine similarity between the query hash code and the hash codes of all resources is then calculated in the hash space. The most relevant resources are then sorted by distance or similarity and returned as the target query result.
[0100] In some other embodiments, in order to improve the accuracy and relevance of the retrieval, the image search intent feature matrix and the text search intent feature matrix may be combined to reorder the retrieval results to improve the accuracy and relevance of the retrieval.
[0101] For example, referring to Figure 3 As shown, when determining the target query result, the following steps S301-S302 may also be included:
[0102] S301 : When the information to be queried is image information to be queried, calculating an image hash code of the image information to be queried based on the image hash function, and determining a target text query result based on the image hash code and the hash code matrix.
[0103] Here, when the query information is an image, an image hash code is calculated based on the image hash function. The target text query result is then determined based on the image hash code and the hash code matrix. The hash code matrix, determined during the previous learning process, contains the semantic association information between the image and the text. By matching it with the image hash code, the relevant text resources can be found.
[0104] S302 : When the information to be queried is text information to be queried, calculating a text hash code of the text information to be queried based on the text hash function, and determining a target image query result based on the text hash code and the hash code matrix.
[0105] Similarly, when the information to be queried is text information, the text hash code of the text information to be queried is calculated based on the text hash function, and then the target image query result is determined based on the text hash code and the hash code matrix, thereby realizing the retrieval from text query to related image resources.
[0106] The MOOCs graphic and text resource cross-modal retrieval method, apparatus, medium, and device provided in the embodiments of the present disclosure improve the accuracy of cross-modal retrieval by introducing the content features of images and texts and combining them with search intent features, and can effectively capture the learners' real search needs. In addition, the use of hash coding matrices for indexing image and text resources not only speeds up the retrieval process but also improves computational efficiency. In this way, more accurate graphic and text resource retrieval can be achieved, effectively improving the accuracy and efficiency of retrieval, reducing computational complexity, meeting the complex and changing learning needs on the MOOCs platform, and improving user experience.
[0107] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0108] Based on the same inventive concept, the embodiment of the present disclosure also provides a MOOCs graphic and text resource cross-modal retrieval device corresponding to the MOOCs graphic and text resource cross-modal retrieval method. Since the principle of solving the problem by the device in the embodiment of the present disclosure is similar to the above-mentioned MOOCs graphic and text resource cross-modal retrieval method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0109] Reference Figure 4 FIG. 4 is a schematic diagram of a MOOCs graphic and text resource cross-modal retrieval device 400 provided in an embodiment of the present disclosure, wherein the device includes:
[0110] Resource acquisition module 401, used to acquire image resource collection and text resource collection of MOOCs platform;
[0111] A matrix construction module 402 is configured to construct an image content feature matrix and an image search intent feature matrix based on the image resource set, and to construct a text content feature matrix and a text search intent feature matrix based on the text resource set;
[0112] A model building module 403 is configured to build a MOOCs resource image-text cross-modal retrieval model based on the image content feature matrix, the image search intent feature matrix, the text content feature matrix, and the text search intent feature matrix;
[0113] Function construction module 404, for determining a hash coding matrix based on the MOOCs resource image-text cross-modal retrieval model, constructing an image hash function based on the image content feature matrix and the hash coding matrix, and constructing a text hash function based on the text content feature matrix and the hash coding matrix;
[0114] The information query module 405 is configured to obtain the information to be queried and determine a target query result based on the information to be queried, the hash coding matrix, the image hash function, and the text hash function.
[0115] In some possible embodiments, the MOOCs resource graphic and text cross-modal retrieval model is expressed as:
[0116] ;
[0117] in, Expressed as a cross-modal retrieval objective function; is expressed as a weight coefficient, and , (i=1,2,…,7), ; Represented as image content feature matrix; Represented as the basis vector of the image content feature matrix; Represents the potential shared content features of image resources and text resources; F is represented as the square root of the sum of the squares of all elements of the matrix; Represented as a text content feature matrix; Represented as the basis vector of the text content feature matrix; Represented as a potentially shared semantic description The projection matrix on the hash space; B is represented as the hash encoding matrix; Represented as an image search intent feature matrix; Represented as the basis vector of the image search intent feature matrix; Representing the potential shared search intent features of image resources and text resources; Represented as a text search intent feature matrix; Represented as the basis vector of the text search intent feature matrix; Represented as potential shared search intent features Projection matrix onto hash space.
[0118] In some possible embodiments, the function construction module 404 is specifically configured to execute:
[0119] Step 1: Random Initialization 、 、 、 、 、 、 、 , B;
[0120] Step 2: Fix other variables in turn and update the single variable according to the following rules:
[0121] The update expression is expressed as:
[0122] ;
[0123] Where I represents the identity matrix; Expressed as The transposed matrix of
[0124] The update expression is expressed as:
[0125] ;
[0126] The update expression is expressed as:
[0127] ;
[0128] The update expression is expressed as:
[0129] ;
[0130] The update expression is expressed as:
[0131] ;
[0132] The update expression is expressed as:
[0133] ;
[0134] The update expression is expressed as:
[0135] ;
[0136] The update expression is expressed as:
[0137] ;
[0138] The update expression of B is expressed as:
[0139] ;
[0140] Step 3: Based on the updated 、 、 、 、 、 、 、 , B. solving the MOOCs resource graphic and text cross-modal retrieval model;
[0141] Step 4: Repeat steps 2 to 3 until the number of iterations reaches a preset value and / or the cross-modal retrieval objective function converges, and determine the hash coding matrix.
[0142] In some possible embodiments, the function construction module 404 is specifically configured to:
[0143] Constructing an initial image hash function based on the image content feature matrix and the hash coding matrix; and constructing an initial text hash function based on the text content feature matrix and the hash coding matrix;
[0144] Constructing a hash target function based on the initial image hash function and the initial text hash function; and solving an image parameter matrix corresponding to the initial image hash function and a text parameter matrix corresponding to the initial text hash function based on the hash target function;
[0145] The image hash function is determined based on the image content feature matrix, the image parameter matrix, and the initial image hash function, and the text hash function is determined based on the text content feature matrix, the text parameter matrix, and the initial text hash function.
[0146] In some possible embodiments, the initial image hash function is expressed as:
[0147] ;
[0148] The initial text hash function is expressed as:
[0149] ;
[0150] The hash objective function is expressed as:
[0151] ;
[0152] in, It is represented as a hash objective function; n represents the number of resources; Represented as an image parameter matrix; Represented as a text parameter matrix; Represented as image content feature matrix The feature vector corresponding to the i-th image resource; Represented as the i-th column of the hash coding matrix; Represented as a text content feature matrix The feature vector corresponding to the i-th text resource.
[0153] In some possible embodiments, the function construction module 404 is specifically configured to:
[0154] Step a: Initialization and ;
[0155] Step b: Update separately by gradient descent method and ,include:
[0156] ;
[0157] ;
[0158] in, Expressed as element-wise multiplication; Represented as the step size of gradient descent;
[0159] Step c: Based on the updated and Solving the hash objective function;
[0160] Step d: Repeat steps b to c until the hash objective function converges and determine and .
[0161] In some possible embodiments, the information to be queried is one of image information to be queried or text information to be queried; the information query module 405 is specifically configured to:
[0162] When the information to be queried is image information to be queried, calculating an image hash code of the image information to be queried based on the image hash function, and determining a target text query result based on the image hash code and the hash code matrix;
[0163] When the information to be queried is text information to be queried, a text hash code of the text information to be queried is calculated based on the text hash function, and a target image query result is determined based on the text hash code and the hash code matrix.
[0164] Based on the same technical concept, the embodiment of the present disclosure also provides a computer device. Figure 5 , which is a schematic diagram of the structure of a computer device 500 provided in an embodiment of the present disclosure, includes a processor 501, a memory 502, and a bus 503. The memory 502 is used to store execution instructions and includes a memory 5021 and an external memory 5022. The memory 5021 is also referred to as internal memory and is used to temporarily store operation data in the processor 501 and data exchanged with an external memory 5022 such as a hard disk. The processor 501 exchanges data with the external memory 5022 through the memory 5021.
[0165] In the embodiment of the present application, the memory 502 is specifically used to store application code for executing the solution of the present application, and the execution is controlled by the processor 501. That is, when the computer device 500 is running, the processor 501 communicates with the memory 502 via the bus 503, so that the processor 501 executes the application code stored in the memory 502, thereby performing the method described in any of the aforementioned embodiments.
[0166] The memory 502 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0167] Processor 501 may be an integrated circuit chip with signal processing capabilities. Such processors may be general-purpose processors, including central processing units (CPUs) and network processors (NPs). They may also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. These processors may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. A general-purpose processor may be a microprocessor or any conventional processor.
[0168] It should be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the computer device 500. In other embodiments of the present application, the computer device 500 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The components shown in the illustrations may be implemented in hardware, software, or a combination of software and hardware.
[0169] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program executes the steps of the method for cross-modal retrieval of MOOCs graphic resources described in the above method embodiment. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0170] The embodiments of the present disclosure further provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the cross-modal retrieval method for MOOCs graphic resources described in the above method embodiment. For details, please refer to the above method embodiment and will not be repeated here.
[0171] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).
[0172] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the several embodiments provided in the present disclosure, it should be understood that the disclosed system and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0173] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0174] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0175] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0176] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The scope of protection of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in the present disclosure, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure shall be subject to the scope of protection of the claims.
Claims
1. A cross-modal retrieval method for MOOCs graphic resources, characterized by: include: Obtain image resource collections and text resource collections from the MOOCs platform; Constructing an image content feature matrix and an image search intent feature matrix based on the image resource set, and constructing a text content feature matrix and a text search intent feature matrix based on the text resource set; Based on the image content feature matrix, the image search intent feature matrix, the text content feature matrix, and the text search intent feature matrix, a MOOCs resource image-text cross-modal retrieval model is constructed; Determine a hash coding matrix based on the MOOCs resource image-text cross-modal retrieval model, construct an image hash function based on the image content feature matrix and the hash coding matrix, and construct a text hash function based on the text content feature matrix and the hash coding matrix; Acquire information to be queried, and determine a target query result based on the information to be queried, the hash coding matrix, the image hash function, and the text hash function; The MOOCs resource graphic and text cross-modal retrieval model is expressed as: ; in, Expressed as a cross-modal retrieval objective function; is expressed as a weight coefficient, and , (i=1,2,…,7), ; Represented as image content feature matrix; Represented as the basis vectors of the image content feature matrix; Represents the potential shared content features of image resources and text resources; F is represented as the square root of the sum of the squares of all elements of the matrix; Represented as a text content feature matrix; Represented as the basis vector of the text content feature matrix; Represented as a potentially shared semantic description The projection matrix on the hash space; B is represented as the hash encoding matrix; Represented as an image search intent feature matrix; Represented as the basis vector of the image search intent feature matrix; Representing the potential shared search intent features of image resources and text resources; Represented as a text search intent feature matrix; Represented as the basis vector of the text search intent feature matrix; Represented as potential shared search intent features Projection matrix onto hash space.
2. The method according to claim 1, characterized in that The determining of the hash coding matrix based on the MOOCs resource graphic and text cross-modal retrieval model includes: Step 1: Random Initialization 、 、 、 、 、 、 、 , B; Step 2: Fix other variables in turn and update the single variable according to the following rules: The update expression is expressed as: ; Where I represents the identity matrix; Expressed as The transposed matrix of The update expression is expressed as: ; The update expression is expressed as: ; The update expression is expressed as: ; The update expression is expressed as: ; The update expression is expressed as: ; The update expression is expressed as: ; The update expression is expressed as: ; The update expression of B is expressed as: ; Step 3: Based on the updated 、 、 、 、 、 、 、 , B. solving the MOOCs resource graphic and text cross-modal retrieval model; Step 4: Repeat steps 2 to 3 until the number of iterations reaches a preset value and / or the cross-modal retrieval objective function converges, and determine the hash coding matrix.
3. The method according to claim 1, characterized in that The constructing of an image hash function based on the image content feature matrix and the hash coding matrix, and the constructing of a text hash function based on the text content feature matrix and the hash coding matrix, include: Constructing an initial image hash function based on the image content feature matrix and the hash coding matrix; and constructing an initial text hash function based on the text content feature matrix and the hash coding matrix; Constructing a hash target function based on the initial image hash function and the initial text hash function; and solving an image parameter matrix corresponding to the initial image hash function and a text parameter matrix corresponding to the initial text hash function based on the hash target function; The image hash function is determined based on the image content feature matrix, the image parameter matrix, and the initial image hash function, and the text hash function is determined based on the text content feature matrix, the text parameter matrix, and the initial text hash function.
4. The method according to claim 3, characterized in that The initial image hash function is expressed as: ; The initial text hash function is expressed as: ; The hash objective function is expressed as: ; in, It is represented as a hash objective function; n represents the number of resources; Represented as an image parameter matrix; Represented as a text parameter matrix; Represented as image content feature matrix The feature vector corresponding to the i-th image resource; Represented as the i-th column of the hash coding matrix; Represented as a text content feature matrix The feature vector corresponding to the i-th text resource.
5. The method according to claim 4, characterized in that The step of respectively solving the image parameter matrix corresponding to the initial image hash function and the text parameter matrix corresponding to the initial text hash function based on the hash target function includes: Step a: Initialization and ; Step b: Update separately by gradient descent method and ,include: ; ; in, Expressed as element-wise multiplication; Represented as the step size of gradient descent; Step c: Based on the updated and Solving the hash objective function; Step d: Repeat steps b to c until the hash objective function converges and determine and .
6. The method according to claim 1, characterized in that The information to be queried is one of image information to be queried or text information to be queried; and determining a target query result based on the information to be queried, the hash coding matrix, the image hash function, and the text hash function includes: When the information to be queried is image information to be queried, calculating an image hash code of the image information to be queried based on the image hash function, and determining a target text query result based on the image hash code and the hash code matrix; When the information to be queried is text information to be queried, a text hash code of the text information to be queried is calculated based on the text hash function, and a target image query result is determined based on the text hash code and the hash code matrix.
7. A cross-modal retrieval device for MOOCs graphic and text resources, characterized by: include: Resource acquisition module, used to obtain image resource collection and text resource collection of MOOCs platform; A matrix construction module, configured to construct an image content feature matrix and an image search intent feature matrix based on the image resource set, and to construct a text content feature matrix and a text search intent feature matrix based on the text resource set; A model construction module is used to construct a MOOCs resource image-text cross-modal retrieval model based on the image content feature matrix, the image search intent feature matrix, the text content feature matrix, and the text search intent feature matrix; A function construction module is used to determine a hash coding matrix based on the MOOCs resource image-text cross-modal retrieval model, construct an image hash function based on the image content feature matrix and the hash coding matrix, and construct a text hash function based on the text content feature matrix and the hash coding matrix; An information query module, configured to obtain information to be queried, and determine a target query result based on the information to be queried, the hash coding matrix, the image hash function, and the text hash function; The MOOCs resource graphic and text cross-modal retrieval model is expressed as: ; in, Expressed as a cross-modal retrieval objective function; is expressed as a weight coefficient, and , (i=1,2,…,7), ; Represented as image content feature matrix; Represented as the basis vectors of the image content feature matrix; Represents the potential shared content features of image resources and text resources; F is represented as the square root of the sum of the squares of all elements of the matrix; Represented as a text content feature matrix; Represented as the basis vector of the text content feature matrix; Represented as a potentially shared semantic description The projection matrix on the hash space; B is represented as the hash encoding matrix; Represented as an image search intent feature matrix; Represented as the basis vector of the image search intent feature matrix; Representing the potential shared search intent features of image resources and text resources; Represented as a text search intent feature matrix; Represented as the basis vector of the text search intent feature matrix; Represented as potential shared search intent features Projection matrix onto hash space.
8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
9. A computer device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Depth unsupervised cross-modal retrieval method for reconstructing Hash based on modal fusion
CN115687571A
Text content and image analysis combined multi-modal content retrieval method and system
CN116204706A