Resource recommendation method and device and storage medium
By performing module detection and semantic vector extraction on documents, we intelligently recommend related resources, which solves the problem of time-consuming and labor-intensive search of resources when processing documents, and improves efficiency.
Patent Information
- Application Number
- CN202510398754.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
When users need to consult related resources when processing documents, the existing technology needs to exit the current document operation interface and search for resources through search engines, resulting in time-consuming and labor-intensive operation and inefficient efficiency.
By obtaining document pictures or documents, dividing them into modules and detecting them, extracting the module's semantic vectors, selecting recommended content from pre-stored reference resources, and providing intelligent recommendation functions.
Intelligently recommend the required resources when users process documents, improve operational efficiency and reduce the time and energy of finding resources.
Smart Images

Figure CN120256728A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to, but are not limited to, the field of artificial intelligence technology, and in particular, to a resource recommendation method, an apparatus, and a storage medium. Background Art
[0002] When processing a document, a user often needs to consult relevant reference materials to help understand the article. At this time, the user usually exits the current document operation interface, opens a browser interface separately, and uses a search engine to perform a query. The user needs to find the required resources from a large number of search results of the search engine, and this process will consume a lot of time and energy of the user, resulting in a relatively low efficiency of the user in processing the document. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in this article. This overview is not intended to limit the scope of protection of the claims.
[0004] Embodiments of the present disclosure provide a resource recommendation method, including: in response to a first operation of a user, obtaining a document picture or receiving a document input by the user; dividing the document picture or the document into one or more modules, and detecting the content of the one or more modules; extracting semantic vectors of at least one module according to the detected content of the module, and selecting one or more reference resources from the semantic vectors of a plurality of pre-stored reference resources as recommended content corresponding to the module.
[0005] Embodiments of the present disclosure further provide a resource recommendation apparatus, including a memory; and a processor connected to the memory, where the memory is used to store instructions, and the processor is configured to execute the steps of the resource recommendation method according to any embodiment of the present disclosure based on the instructions stored in the memory.
[0006] Embodiments of the present disclosure further provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the resource recommendation method according to any embodiment of the present disclosure is implemented.
[0007] Embodiments of the present disclosure further provide a program product, including instructions, and when the computer program product is executed by a computer, the instructions execute the resource recommendation method according to any embodiment of the present disclosure.
[0008] An embodiment of the present disclosure also provides a resource recommendation device, including a data acquisition module, a layout analysis module, and a recommendation module, where: the data acquisition module is configured to obtain a document picture or receive a document input by a user in response to a first operation of the user; the layout analysis module is configured to divide the document picture or the document into one or more modules and detect the content of the one or more modules; the recommendation module is configured to extract semantic vectors of at least one module according to the content of the detected module, and select one or more of the reference resources from the semantic vectors of a plurality of pre-stored reference resources as the recommended content corresponding to the module.
[0009] The resource recommendation method, device, and storage medium according to the embodiments of the present disclosure obtain a document picture or receive a document input by a user in response to a first operation of the user; divide the document picture or the document into one or more modules and detect the content of the one or more modules; extract semantic vectors of at least one module according to the content of the detected module, and select one or more of the reference resources from the semantic vectors of a plurality of pre-stored reference resources as the recommended content corresponding to the module, which can intelligently recommend the required reference resources to the user when the user processes the document, facilitate the user operation, and improve the efficiency of the user processing the document.
[0010] Other features and advantages of the present disclosure will be described in the following specification, and some of them will become obvious from the specification, or be understood by implementing the present disclosure. Other advantages of the present disclosure can be realized and obtained by the solutions described in the specification and the drawings. Description of the Drawings
[0011] The drawings are used to provide an understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0012] Figure 1 It is a schematic flowchart of a resource recommendation method provided for an exemplary embodiment of the present disclosure;
[0013] Figure 2 It is a schematic diagram of an icon interface of a resource recommendation method provided for an exemplary embodiment of the present disclosure;
[0014] Figure 3 It is a schematic diagram of a function list interface provided for an exemplary embodiment of the present disclosure;
[0015] Figure 4 It is a schematic flowchart of a layout analysis provided for an exemplary embodiment of the present disclosure;
[0016] Figure 5Schematic diagram of the layout analysis result of a Chinese content provided by an exemplary embodiment of the present disclosure;
[0017] Figure 6 Schematic diagram of the layout analysis result of an English content provided by an exemplary embodiment of the present disclosure;
[0018] Figure 7A For Figure 2 the document operation interface shown, schematic diagram of the intelligent recommendation result that appears after clicking the intelligent assistant icon;
[0019] Figure 7B For Figure 7A specific content display diagram of the intelligent recommendation result in
[0020] Figure 8 Schematic diagram of the structure of a resource recommendation device provided by an exemplary embodiment of the present disclosure;
[0021] Figure 9 Schematic diagram of the structure of another resource recommendation device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners
[0022] The present disclosure describes multiple embodiments, but the description is exemplary rather than restrictive, and it is obvious to those of ordinary skill in the art that there can be more embodiments and implementation solutions within the scope of the embodiments described in the present disclosure. Although many possible feature combinations are shown in the drawings and discussed in the detailed implementation manners, many other combination ways of the disclosed features are also possible. Unless specifically restricted, any feature or element of any embodiment can be combined with any other feature or element in any other embodiment, or can replace any other feature or element in any other embodiment.
[0023] The present disclosure includes and contemplates combinations with features and elements known to those of ordinary skill in the art. The disclosed embodiments, features, and elements of the present disclosure can also be combined with any conventional features or elements to form unique invention solutions defined by the claims. Any feature or element of any embodiment can also be combined with features or elements from other invention solutions to form another unique invention solution defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in the present disclosure can be implemented alone or in any suitable combination. Therefore, the embodiments are not subject to other restrictions except those made according to the appended claims and their equivalent replacements. In addition, various modifications and changes can be made within the scope of protection of the appended claims.
[0024] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not depend on the specific order of the steps described herein, the method or process should not be limited to the specific order of steps described. As those of ordinary skill in the art will understand, other step orders are possible. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation on the claims. In addition, the claims directed to the method and / or process should not be limited to performing their steps in the order written, as those skilled in the art can readily understand that these orders can vary and still remain within the spirit and scope of the embodiments of the present disclosure.
[0025] As Figure 1 shown, embodiments of the present disclosure provide a resource recommendation method, including:
[0026] Step 101, in response to a first operation of a user, obtain a document picture or receive a document input by the user;
[0027] Step 102, divide the document picture or the document into one or more modules, and detect the content of the one or more modules;
[0028] Step 103, extract semantic vectors of at least one module according to the content of the detected module, and select one or more reference resources from the semantic vectors of a plurality of pre-stored reference resources as recommended content corresponding to the module according to the extracted semantic vectors.
[0029] The resource recommendation method of the embodiments of the present disclosure, by responding to a first operation of a user, obtaining a document picture or receiving a document input by the user, dividing the document picture or the document into one or more modules, detecting the content of the one or more modules, extracting semantic vectors of at least one module according to the content of the detected module, and selecting one or more reference resources from the semantic vectors of a plurality of pre-stored reference resources as recommended content corresponding to the module, can intelligently recommend required reference resources to the user when the user reads or edits a document, facilitate the user's operation, and improve the efficiency of the user in processing the document.
[0030] The resource recommendation method of the embodiments of the present disclosure can be implemented as a first icon (hereinafter referred to as the intelligent assistant icon) suspended on the document operation interface as Figure 2 shown. The document operation interface is the interface provided by document operation software (such as a document reader and / or a document editor). Among them, the shape, size, position, and transparency of the first icon can be set as needed, and the present disclosure places no restrictions thereon.
[0031] As Figure 2As shown, when the user reads or edits the document content on the document operation interface, the first icon can be displayed in real time on the sidebar of the document operation interface. Exemplarily, when the user clicks Figure 2 the intelligent assistant icon in, the function list interface as shown in Figure 3 is popped up. In the embodiments of the present disclosure, the function list interface may include, but is not limited to, a screenshot control trigger button, an import control trigger button, a layout analysis control trigger button, a recommendation control trigger button, a storage control trigger button, etc. Exemplarily, as Figure 3 shown, the screenshot control trigger button can be implemented in the form of a screenshot button, the import control trigger button can be implemented in the form of an import button, the layout analysis control trigger button can be implemented in the form of a layout analysis slide switch button, the recommendation button can be implemented in the form of an intelligent recommendation slide switch button and a personalized recommendation slide switch button, and the storage control trigger button can be implemented in the form of a structure storage slide switch button. Each button is used to display one or more functions of the intelligent assistant. However, the present disclosure does not limit this. The structure of the function list interface can be set according to the functions provided by the intelligent assistant. For example, in some other examples, only any one of the intelligent recommendation function and the personalized recommendation function can be set.
[0032] In the embodiments of the present disclosure, the functions provided by the intelligent assistant include, but are not limited to, screenshot (i.e., screen capture), import, layout analysis, intelligent recommendation, personalized recommendation, structure storage, and other functions. Among them, the screenshot function is used for the user to take a screenshot of the current display screen (the captured picture is used as a document picture); the import function is used for the user to input one or more documents; the layout analysis function is used to divide the received document picture or document into one or more modules and detect the content of each module; the intelligent recommendation function is used to recommend one or more relevant reference resources according to the content of the detected module; the personalized recommendation is used to recommend relevant reference resources that the user is more interested in according to the content of the detected module and the user's interest weight; the structure storage is used to save the document picture or document together with the recommended reference resources.
[0033] In some exemplary embodiments, the first operation may be a trigger operation on the screenshot control or an import control. Among them, the trigger operation on the screenshot control may include at least one of the following:
[0034] After opening the document, click the screenshot control trigger button on the function list interface;
[0035] Open the document;
[0036] After the position of the scroll bar on the document operation interface changes, the stop duration reaches the preset stop duration threshold.
[0037] Among them, the triggering operations on the import control include at least one of the following:
[0038] Click the trigger button of the import control on the function list interface;
[0039] Drag the document into the first icon, where the function list interface is the interface presented after operating on the first icon.
[0040] In the embodiments of the present disclosure, after the layout analysis function is turned on, as long as the user opens a document or scrolls the document operation interface during the process of viewing / editing the document, within a preset duration after the scrolling stops, the intelligent assistant will automatically take a screenshot of the current document operation interface (at this time, the screenshot control takes a screenshot of the entire document operation interface), perform layout analysis based on the obtained screenshot image, and display the layout analysis result on the document operation interface.
[0041] In the embodiments of the present disclosure, after the recommendation (intelligent recommendation or personalized recommendation) function is turned on, as long as the user opens a document or scrolls the document operation interface during the process of viewing / editing the document, within a preset duration after the scrolling stops, the intelligent assistant will automatically take a screenshot of the current document operation interface (at this time, the screenshot control takes a screenshot of the entire document operation interface), perform layout analysis and recommendation based on the obtained screenshot image. In this way, when the user opens a document, relevant recommended content will appear, and as the user scrolls the page, the relevant recommended content will also be automatically updated. In the embodiments of the present disclosure, it can be preset that the intelligent assistant only displays the recommendation result on the document operation interface and does not display the layout analysis result, or displays both the layout analysis result and the recommendation result at the same time. For non-developer users, usually, only the recommendation result needs to be displayed, and the layout analysis result does not need to be displayed.
[0042] In the embodiments of the present disclosure, in step 101, when obtaining the document image, the first operation of the user can be to open the document, or click the trigger button of the screenshot control on the function list interface, or stop after the position of the scroll bar of the document operation interface changes and the stop duration reaches the preset stop duration threshold (for example, stop scrolling after scrolling the mouse wheel on the document operation interface for the preset stop duration threshold, or stop dragging after dragging the scroll bar on the document operation interface for the preset stop duration threshold). In the embodiments of the present disclosure, the preset stop duration threshold can be set as needed. For example, the preset stop duration threshold can be set to 1 second or 2 seconds. However, the present disclosure does not limit this. Usually, the process of the intelligent assistant performing layout analysis and intelligent recommendation (or personalized recommendation) is within 1 second. Therefore, when the user scrolls the display range of the document while viewing or editing the document, when the scrolling stops for about 2 seconds to 3 seconds, the intelligent assistant can display the relevant recommended content.
[0043] Exemplarily, when the user clicks the trigger button of the screenshot control, the mouse pointer will become a crosshair, and the user can drag the mouse to select the area to be captured. The captured area can be of any shape. For example, it can be the visible area of the current document operation interface (i.e., the document area currently visible to the user) or a specific part within the visible area of the current document operation interface. When the user scrolls the mouse wheel on the document operation interface, the intelligent assistant will re-acquire the document picture and perform layout analysis and recommendation based on the acquired document picture. In this way, as the user scrolls the page, the relevant recommended content will be automatically updated.
[0044] In the embodiments of the present disclosure, in step 101, when receiving the document input by the user, the first operation of the user can be to click the trigger button of the import control on the function list interface or drag the document into the first icon. Exemplarily, when the user clicks the trigger button of the import control, a file browser window is launched, the user selects the document to be imported, and then clicks OK to complete the file import process. In some other examples, the user can also trigger the operation of the intelligent assistant to receive the document input by the user by directly dragging the document (including documents in pure picture format or documents in any other format) into the intelligent assistant icon.
[0045] However, the present disclosure does not limit this. In some other examples, the specific implementation manner of the first operation can be adjusted as needed. For example, the first operation of the user can also be to double-click the intelligent assistant icon, etc.
[0046] In the embodiments of the present disclosure, the intelligent assistant provides the following two input methods:
[0047] (1) Real-time screenshot: The user can trigger the button of the screenshot control on the function list interface of the intelligent assistant to capture pictures in real time. After the user successfully takes a screenshot, the intelligent assistant automatically performs layout analysis and recommendation on the user's screenshot.
[0048] (2) Document input: The user can drag the document into the icon of the intelligent assistant or open the corresponding document by triggering the button of the import control in the function list interface of the intelligent assistant. After the document is dragged into the icon of the intelligent assistant or imported by triggering the button of the import control on the function list interface, the intelligent assistant automatically performs layout analysis and recommendation on the document content.
[0049] In the embodiments of the present disclosure, in step 101, the acquired document picture can be the user's current document operation interface (for example, the document reading interface or the document editing interface). However, the present disclosure does not limit this.
[0050] In some other examples, the acquired document picture can also be a picture converted from the document. For example, when the user's document includes multiple pages, the acquired document pictures include multiple pictures, and each document picture corresponds to one page of the document.
[0051] In some exemplary embodiments, in step 101, the obtained document picture may be determined according to the upper and lower boundaries of the visible area of the document operation software, and / or according to the user's screenshot interface.
[0052] In the embodiments of the present disclosure, when determining the document picture according to the upper and lower boundaries of the visible area of the document operation software, the document picture is all the visible areas of the document operation software (i.e., the entire document operation interface); when using the user's screenshot interface as the document picture, the document picture is all the visible areas of the document operation software, or a partial area of all the visible areas of the document operation software (i.e., a partial document operation interface), and the present disclosure does not limit this.
[0053] In some exemplary embodiments, the method further includes: in response to a second operation of the user, obtaining the upper and lower boundaries of the visible area of the document operation interface, where the second operation includes: opening a document, and / or changing the position of the scroll bar of the document operation interface and then stopping with a stop duration reaching a preset stop duration threshold.
[0054] In the embodiments of the present disclosure, when the user opens a document or scrolls the mouse wheel (or drags the scroll bar) on the document operation interface, the intelligent assistant can communicate with the document operation software using an application programming interface (API) or a plug-in to obtain the current state of the document operation software (including information such as the visible area, the scrolling state of the scroll bar, and the cursor position).
[0055] Exemplarily, taking a document editor as an example, the upper and lower boundaries of the visible area in the current document editor can be obtained through JavaScript programming. In the embodiments of the present disclosure, the upper boundary of the visible area is the top position of the user's viewport (i.e., the screen area where the document operation interface is displayed) (i.e., the vertical coordinate of the upper edge of the current screen), and the value of the upper boundary will change with the movement of the scroll bar position. For example, if the top position of the current screen area is 100 pixels away from the beginning of the document, then the upper boundary of the visible area is 100 pixels. Similarly, the lower boundary of the visible area is the upper boundary of the visible area plus the height of the document editor, representing the bottom position of the user's viewport. For example, assuming the height of the document editor is 400 pixels and the upper boundary of the visible area is 100 pixels, then the lower boundary of the visible area is 100 + 400 = 500 pixels. After obtaining the upper and lower boundaries of the visible area, the area within the range of the upper and lower boundaries of the visible area is the document area (i.e., the document operation interface) that the user is browsing.
[0056] In some exemplary embodiments, in step 101, the document input by the user may be in any of the following formats: pure picture document, pure text document, Word document, PDF document, PPT document, CAJ document, etc. However, the present disclosure does not limit this. The document input by the user may also be a document in any other format other than the above formats. For example, the document input by the user may also be in formats such as Excel and HTML.
[0057] It should be noted that the resource recommendation method of the embodiments of the present disclosure can also be used during the user's web browsing. For example, when the user is browsing a web page, if the user clicks on the intelligent assistant icon, the intelligent assistant performs layout analysis and detection on the web page being browsed by the user and automatically recommends relevant resource content.
[0058] When using the intelligent assistant of the embodiments of the present disclosure, the layout analysis function and the intelligent recommendation function can be default automatically turned on. After the intelligent assistant obtains a document picture or receives a document input by the user, it automatically sends the document picture or the document into the subsequent layout analysis and intelligent recommendation process. After obtaining the recommended content, it then feeds back the recommended content to the user.
[0059] In the embodiments of the present disclosure, before the intelligent assistant executes the layout analysis function, the document input by the user can be converted into one or more pictures, so that no matter which input method the user uses (screenshot or import), the input of the layout analysis module of the intelligent assistant is a picture. However, the present disclosure does not limit this. In some other examples, the layout analysis module of the intelligent assistant can also directly perform layout analysis on the document input by the user.
[0060] In some exemplary embodiments, in step 102, one or more modules include at least one of the following types: picture, table, formula, text, explanatory text, title, reference.
[0061] In some exemplary embodiments, the method further includes: in response to an operation of the user clicking on the layout analysis control trigger button on the function list interface, turning on the layout analysis function (i.e., dividing the document picture or the document into one or more modules and detecting the content of the one or more modules).
[0062] Figure 4 This is a schematic diagram of an exemplary layout analysis process of the present disclosure. As Figure 4As shown, when performing the layout analysis function, first obtain the document image to be analyzed. This document image can be the document image converted from the document input by the user or the real-time screenshot of the user (the intelligent assistant can select the corresponding document operation software according to the document format and convert the document input by the user into a document image). Then perform layout detection on the document image. The detection categories include figure, table, formula, text, title, reference, etc. Example detection results are shown in Figure 5 or Figure 6 as shown Figure 5 is a schematic diagram of an example of the layout analysis result of a Chinese content Figure 6 is a schematic diagram of an example of the layout analysis result of an English content. From Figure 5 and Figure 6 it can be seen that the detection effect can meet the subsequent recognition requirements. Then perform content recognition on each detected module. When the detection result is a table, perform table content recognition and save the table after cutting; when the detection result is a formula, recognize the formula; when the detection result is text, since there may be formulas mixed in the text, further detect whether there are formulas in the text. If there are formulas, perform formula recognition on the formula part and text recognition on the other parts; if there are no formulas, only perform text recognition; when the detection result is a picture, save the picture after cutting (in some other examples, it is also possible to recognize the picture content). Finally, integrate the above content into a standard format according to a certain post-processing strategy, such as uniformly generating Markdown format (Markdown is a lightweight markup language that supports figures, tables, mathematical formulas, etc.), which is convenient for later editing and modification.
[0063] In the embodiments of the present disclosure, the intelligent assistant can implement the layout analysis (including layout detection, text recognition, chart recognition, formula recognition, etc.) function based on a convolutional neural network model or a Transformer neural network model.
[0064] In some exemplary embodiments, in step 102, detecting the content of one or more modules includes:
[0065] Select at least one module;
[0066] Determine the corresponding content recognition model according to the type of the selected module;
[0067] Input the selected module into the corresponding content recognition model to obtain the content of the module.
[0068] Exemplarily, the above content recognition model may include, but is not limited to, a text recognition model, an image recognition model, a table recognition model, a formula recognition model, etc. Among them, the text recognition model, the image recognition model, the table recognition model, and the formula recognition model may all be convolutional neural network models or Transformer neural network models. However, the present disclosure does not limit this.
[0069] In the embodiments of the present disclosure, for functions such as layout detection, text recognition, chart recognition, formula recognition, etc., corresponding layout detection models, text recognition models, image recognition models, table recognition models, formula recognition models, etc. can be pre-trained respectively. The input image is divided into one or more modules using the layout detection model, the content of the input table is detected using the table recognition model, the content of the input image is detected using the image recognition model, the content of the input formula is detected using the formula recognition model, and the content of the input text is detected using the text recognition model.
[0070] In some exemplary embodiments, in step 102, dividing the document image or document into one or more modules includes:
[0071] Inputting the document image or document into the layout detection model to obtain the division result of the document image or document. Among them, when the number of pages of the document image or the number of pages of the document is less than or equal to the preset page number threshold, the layout detection model is a local model; when the number of pages of the document image or the number of pages of the document is greater than the preset page number threshold, the layout detection model is a cloud model.
[0072] The intelligent assistant in the embodiments of the present disclosure includes two computing capabilities, one is local computing and the other is cloud computing. For local computing, considering the limitation of computing resources, a lightweight local model can be deployed to only perform layout analysis and subsequent calculations on screenshot images and documents with fewer pages. For longer and more complex documents, it automatically switches to cloud computing. Since cloud servers have strong computing capabilities, full-parameter large models can be deployed, which can quickly process a large amount of data. Through this hybrid solution, while ensuring the user experience, the powerful computing capabilities and update convenience of the cloud can be utilized.
[0073] In some exemplary embodiments, the method further includes: in response to the user's operation of clicking the button of the recommendation control on the function list interface, displaying the recommended content of the corresponding module in a pop-up window or a collapsible panel.
[0074] In the embodiments of the present disclosure, the recommendation function can be default automatically turned on, or turned on after the user clicks the button of the recommendation control on the function list interface.
[0075] Exemplarily, assuming the recommendation function is turned on, when the user clicks Figure 2When the smart assistant icon in Figure 7A is shown, as shown in Figure 7B , the smart assistant can display the content it recommends in the form of a pop-up window ( Figure 7A is the specific recommended result display diagram in
[0076] . The pop-up window generally floats on the side of the page, and the transparency of the pop-up window is adjustable. By adjusting the transparency, it can be made not to block the main text. The user can drag the pop-up window to any position. As the user scrolls the page, the relevant recommended content in the pop-up window will be automatically updated. When the user interacts with the pop-up window (for example, clicks on any area of the pop-up window), the pop-up window can automatically adjust its size according to the size of the current recommended content. For example, when there is more current recommended content, the size of the pop-up window is automatically enlarged. Or, the user can use the mouse to adjust the window size. For example, move the mouse cursor to the edge of the window, the cursor shape becomes a double-headed arrow, the user clicks and holds the left mouse button, and at the same time, moves the mouse to change the size of the window. After adjusting the window size, release the left mouse button.
[0077] In some exemplary embodiments, in step 103, extracting at least one semantic vector of the module according to the detected content of the module includes:
[0078] Extracting the first information of the module according to the detected content of the module;
[0079] Inputting the extracted first information into the first pre-trained model to obtain the semantic vector of the module.
[0080] In the embodiments of the present disclosure, the first information may be at least one of the following information: abstract, keyword, title, content, etc. However, the present disclosure does not limit this.
[0081] In the embodiments of the present disclosure, the first pre-trained model may be a natural language processing model such as BERT (Bidirectional Encoder Representations from Transformers) or Transformer. However, the present disclosure does not limit this.
[0082] In some exemplary embodiments, in step 103, selecting one or more reference resources from the semantic vectors of a plurality of pre-stored reference resources as the recommended content corresponding to the module includes any one of the following:
[0083] Calculate the similarity between the extracted semantic vector and the semantic vectors of multiple pre-stored reference resources, and select one or more reference resources as the recommended content for the corresponding module according to the calculation results of the similarity between the extracted semantic vector and the semantic vectors of the multiple reference resources;
[0084] Input the extracted semantic vector into a trained recommendation model to obtain the reference resources output by the recommendation model, and use the reference resources output by the recommendation model as the recommended content for the corresponding module, where the recommendation model is pre-trained using the semantic vectors of multiple pre-stored reference resources;
[0085] Calculate the distance between the extracted semantic vector and multiple class center vectors, and select one or more reference resources as the recommended content for the corresponding module from the classes in which the calculated distance is lower than the preset distance threshold, where the multiple class center vectors are obtained by performing clustering analysis on the semantic vectors of multiple pre-stored reference resources.
[0086] In the embodiments of the present disclosure, the method for selecting one or more reference resources from the semantic vectors of multiple pre-stored reference resources as the recommended content for the corresponding module according to the extracted semantic vector is not limited to the above methods, and users can use other selection methods according to their needs, and the present disclosure does not limit this.
[0087] Exemplarily, when obtaining the recommended content for the corresponding module by the method based on the pre-trained recommendation model, a deep learning model can be used as the recommendation model, and a corresponding category label is set for each reference resource. The semantic vectors and category labels of the multiple reference resources are used as training data to train the recommendation model. After the training is completed, the extracted semantic vector is input into the recommendation model, and then the category prediction can be directly performed, and the reference resources corresponding to the predicted category are used as the recommended content for the corresponding module.
[0088] Exemplarily, when obtaining the recommended content for the corresponding module by the clustering-based method, the semantic vectors of the multiple reference resources can be clustered using a clustering method such as K-Means or any other arbitrary clustering method, the semantic vectors of the multiple reference resources are divided into multiple different classes, and then the semantic vector extracted from the current module is compared with each class center vector to obtain the distance between the semantic vector extracted from the current module and each class center vector. The class with the closest distance is selected as the class to be recommended, and the content to be recommended is flexibly selected (such as randomly selecting from this class, etc.) from the class to be recommended.
[0089] Exemplarily, when selecting one or more reference resources as the recommended content for the corresponding module according to the calculation results of the similarity between the extracted semantic vector and the semantic vectors of the multiple reference resources, a database can be established in advance, and the semantic vectors and recommendation information of the multiple reference resources are stored in the database.
[0090] In some exemplary embodiments, before step 103, the method further includes:
[0091] Pre - establish a database that stores multiple reference resources. Each reference resource includes second information for recommended display and a semantic vector for similarity comparison.
[0092] The intelligent recommendation function of the embodiments of the present disclosure relies on a pre - established database, which can select various different data storage methods, such as MySQL (a relational database management system), MongoDB (a database based on distributed file storage), FAISS (an open - source vector database), etc. The present disclosure places no restrictions on this.
[0093] In the embodiments of the present disclosure, the second information corresponding to each reference resource is used as the content for intelligent recommendation or personalized recommendation, and the semantic vector is used when comparing the similarity between the detected module and the pre - stored reference resources.
[0094] In the embodiments of the present disclosure, the second information may be at least one of the following information: abstract, title, content, website URL, etc. However, the present disclosure places no restrictions on this.
[0095] The intelligent recommendation function of the embodiments of the present disclosure calculates the similarity between the content of one or more modules and multiple pre - stored reference resources, and dynamically recommends reference resources (such as reference documents, data, or context information, etc.) related to the current document or document picture according to the similarity calculation result. The intelligent recommendation is not only based on the text content but also based on the visual elements (such as tables, images, etc.) identified during the layout analysis process, ensuring that users can seamlessly obtain additional information when processing documents.
[0096] The intelligent assistant can perform real - time analysis based on the content that the user is currently viewing or editing. When the user is processing a certain paragraph or chart in a document, the intelligent assistant can recommend relevant reference materials based on the content such as the text and charts to help the user better analyze and understand the document content.
[0097] Such as Figure 7A and Figure 7B As shown, taking the example of assisting users in reading legal provisions, the intelligent assistant can automatically analyze the content of legal documents and give reference content such as cases and legal articles related to the legal provisions to help users understand and analyze the current reading content.
[0098] When applying intelligent assistants to the legal field, the first pre-trained model can be large language models such as Legal-BERT (a BERT model suitable for legal document classification, legal entity recognition, legal question answering, etc. in the legal field), LexGlove (a word embedding model that can generate high-quality word vectors), etc. These large language models are trained through a large number of legal documents and are more suitable for dealing with legal-related semantic matching problems and can better understand legal terms and expressions.
[0099] In this example, a large number of legal cases (public legal databases, court judgments, regulatory documents, etc.) can be collected and stored in a pre-established database after processing.
[0100] Exemplarily, the process of establishing the database may include the following steps:
[0101] (I) Analyze legal cases and extract the core information of the cases, including case names, judgment summaries, judgment contents, etc.
[0102] (II) Convert the case text into semantic vectors through the first pre-trained model (such as Legal-BERT). The input format of the model can be: "[CASE NAME] important judgment case title [SUMMARY] judgment summary [CONTENT] judgment content...", and the model will generate a high-dimensional vector representing the semantic features of the input text.
[0103] (III) Build an efficient query database. This query database stores the semantic vectors of each case and also stores the metadata (i.e., the second information) of each case, including: case ID (unique identifier), case name, judgment summary, and other information.
[0104] In some exemplary embodiments, one or more reference resources are selected as the recommended content for the corresponding module according to the similarity calculation result between the extracted semantic vector and the semantic vectors of multiple reference resources, including:
[0105] Select the reference resources with similarity calculation results higher than the preset similarity threshold as the recommended content for the corresponding module.
[0106] Exemplarily, the reference resource with the largest similarity calculation result can be selected as the recommended content for the corresponding module. However, the present disclosure does not limit this.
[0107] Still taking the example of assisting users in reading legal provisions, when a legal provision is obtained through layout analysis, the model takes this legal provision as input, converts it into a semantic vector, then uses a similarity calculation method to calculate the similarity between the current semantic vector and the semantic vectors pre-stored in the database, and returns the most similar semantic vector. Based on this result, the legal cases most relevant to the current legal provision are obtained (such as returning case IDs, names, judgment abstracts, etc.), and finally these contents are displayed in the reference resources recommended by the intelligent assistant.
[0108] In some exemplary embodiments, the method further includes:
[0109] Display the recommended content of the corresponding module in a pop-up window, wherein the position of the pop-up window is movable, the size is modifiable, and the transparency is settable.
[0110] In the embodiments of the present disclosure, the recommended content can be displayed using a pop-up window. Among them, the pop-up window can be a floating-layer pop-up window, a prompt-box pop-up window, a card pop-up window, etc. Exemplarily, the recommended content can be displayed in the form of a prompt-box pop-up window at the edge of the page without covering the content being viewed by the user (semi-transparency display can be selected).
[0111] In some exemplary embodiments, the method further includes:
[0112] Determine the position of the detected module between the upper and lower boundaries of the visible area of the document operation interface, and determine the display position of the pop-up window according to the position of the module.
[0113] In the embodiments of the present disclosure, the display position of the pop-up window can be arranged as much as possible around the module to be recommended, facilitating the user to read and view.
[0114] In some exemplary embodiments, the method further includes at least one of the following:
[0115] In response to the user's drag operation on the pop-up window, move the position of the pop-up window;
[0116] In response to detecting that the size of the document content covered by the pop-up window exceeds a preset area threshold, automatically move the position of the pop-up window according to the document content.
[0117] In the embodiments of the present disclosure, the display position of the pop-up window can be manually adjusted by the user, or automatically adjusted by the intelligent assistant when it detects that the size of the document content covered by the pop-up window exceeds the preset area threshold. When the intelligent assistant automatically adjusts the display position of the pop-up window, it can try to display the pop-up window around the module to be recommended and try to ensure that the content of the pop-up window does not block the document content, for the convenience of the user to view.
[0118] In some exemplary embodiments, the method further includes at least one of the following:
[0119] In response to a user's operation of resizing a pop-up window, resize the pop-up window;
[0120] In response to detecting a user's interaction operation on the pop-up window, automatically resize the pop-up window according to the size of the recommended content.
[0121] In the embodiments of the present disclosure, the size of the pop-up window can be manually adjusted by the user, or automatically adjusted by the intelligent assistant according to the size of the current recommended content when the intelligent assistant detects the user's interaction operation on the pop-up window. In the embodiments of the present disclosure, the user's interaction operation on the pop-up window can be an operation of the user clicking on the pop-up window interface or a pull-down operation of the user on the pop-up window interface, etc. When the intelligent assistant automatically adjusts the size of the pop-up window, if there is more recommended content, the intelligent assistant can try to use a pop-up window with a relatively large size for display to facilitate the user's viewing.
[0122] In some exemplary embodiments, the method further includes at least one of the following:
[0123] In response to a user's operation of setting the transparency parameter of the pop-up window, adjust the transparency of the pop-up window;
[0124] In response to detecting that the size of the document content covered by the pop-up window exceeds a preset area threshold, automatically adjust the transparency of the pop-up window.
[0125] In the embodiments of the present disclosure, the transparency of the pop-up window can be manually adjusted by the user, or automatically adjusted by the intelligent assistant according to the size of the document content currently covered by the pop-up window. When the intelligent assistant automatically adjusts the transparency of the pop-up window, if there is more document content covered by the pop-up window, the intelligent assistant can try to increase the transparency of the pop-up window to facilitate the user's viewing.
[0126] In some other exemplary embodiments, the method further includes:
[0127] Display the recommended content corresponding to the module in a collapsible panel.
[0128] In the embodiments of the present disclosure, the recommended content can be displayed in a collapsible panel. The collapsible panel can be dragged to the sidebar or the bottom, and the user can choose to expand it for viewing instead of automatically displaying it.
[0129] In some other exemplary embodiments, the method further includes:
[0130] In response to a user's third operation, obtain and display the recommended content corresponding to the module where the intelligent assistant icon is located.
[0131] In this embodiment, the third operation can be dragging the intelligent assistant icon to a certain position in the document operation interface (i.e., the place where the user hopes to obtain the recommended content). At this time, the display of the recommended content is triggered by the third operation.
[0132] In some exemplary embodiments, selecting one or more reference resources as the recommended content for the corresponding module according to the similarity calculation result between the extracted semantic vector and the semantic vectors of multiple reference resources includes:
[0133] Obtaining the user's interest field and the domain feature vectors of multiple reference resources;
[0134] Calculating the domain weight of each reference resource according to the user's interest field and the domain feature vectors of multiple reference resources;
[0135] Calculating the weighted similarity calculation result of each reference resource according to the similarity calculation result and the domain weight of each reference resource;
[0136] Selecting the reference resources with the weighted similarity calculation result higher than the preset similarity threshold as the recommended content for the corresponding module.
[0137] Still taking the example of assisting the user in reading legal provisions, when constructing the database, there are multiple relevant cases (i.e., the above-mentioned reference resources) involving different legal provisions stored in the database, and a multi-label encoding method can be used to identify the legal fields to which each case belongs. Exemplarily, each legal case can be identified by a multi-label feature vector for the legal fields it involves. For example, assuming there are three legal fields in total: criminal law, civil law, and commercial law, if a certain case involves criminal law and commercial law, its domain feature vector can be expressed as: domain feature vector = [1, 0, 1], indicating that this case involves both criminal law and commercial law.
[0138] In addition, an interest weight vector can be constructed for the user according to the historical record or user portrait to identify the user's interest field, that is, the user's preference for different legal fields. For example, for the aforementioned three legal fields: criminal law, civil law, and commercial law, the user's interest weight vector can be expressed as: [0.8, 0.2, 0.5], indicating that the user has a high interest in criminal law, a low interest in civil law, and a medium interest in commercial law.
[0139] In some exemplary embodiments, calculating the domain weight of each reference resource according to the user's interest field and the domain feature vectors of multiple reference resources includes any one of the following:
[0140] Calculating the domain weight of each reference resource as the vector inner product between the domain feature vector of the reference resource and the user interest weight, and taking the calculated vector inner product as the domain weight of each reference resource;
[0141] Calculating the similarity between the domain feature vector of each reference resource and the user interest weight, and taking the similarity calculation result between the domain feature vector of each reference resource and the user interest weight as the domain weight of each reference resource.
[0142] In the embodiments of the present disclosure, the vector inner product between the domain feature vectors of each reference resource and the user interest weights can be used as the domain weights of each reference resource. However, the present disclosure does not limit this. In some other examples, methods such as cosine similarity and Euclidean distance can also be used to calculate the similarity or distance between the domain feature vectors of cases and the user interest weights as the domain weights of each case. Here, taking the method of calculating the vector inner product as an example, assuming that A is the domain feature vector of a case and B is the user interest weight vector, the calculated inner product A·B is the domain weight of this case.
[0143]
[0144] Among them, n represents the dimension of the two vectors. When calculating the inner product, first multiply the corresponding elements of the two vectors, and then add the products of each element.
[0145] For each case, calculate the inner product of the user interest weight vector and the domain feature vector of this case. For example, assuming that the user queries the provisions of "Criminal Law", for the aforementioned three legal fields: Criminal Law, Civil Law, and Commercial Law, the interest weights are [0.8, 0.2, 0.5]. In the database, the domain feature vector of case A is [1, 0, 1] (Criminal Law and Commercial Law), and the domain feature vector of case B is [0, 1, 1] (Civil Law and Commercial Law). Then the domain weight of case A = (0.8×1)+(0.2×0)+(0.5×1) = 0.8 + 0 + 0.5 = 1.3, and the domain weight of case B = (0×1)+(0.2×1)+(0.5×1) = 0 + 0.2 + 0.5 = 0.7.
[0146] When calculating the domain weights of each reference resource by the cosine similarity calculation method, assuming that A is the domain feature vector of a case and B is the user interest weight vector, the calculated cosine similarity cos(A,B) is the domain weight of this case:
[0147]
[0148] For the above case A, the calculated cosine similarity is Similarly, the calculated cosine similarity of case B cos(A,B) = 0.7 / 1.36 = 0.51.
[0149] After calculating the domain weights of all cases, multiply the domain weight of each case by the corresponding text similarity calculation result to obtain the weighted similarity calculation result. The cases can be sorted according to the weighted similarity calculation result, and the case with the highest weighted similarity calculation result is selected for recommendation.
[0150] In the embodiments of the present disclosure, as the user's usage habits and preferences change, the user's interest weight vector can be adjusted according to the user's feedback and behavior. The recommended results finally presented by the intelligent assistant comprehensively consider the user's usage habits and preferences as well as the text similarity calculation results, rather than simply the text similarity calculation results.
[0151] In some exemplary embodiments, the method further includes:
[0152] In response to the user's fourth operation, the content of one or more modules is stored in a preset standard format.
[0153] In the embodiments of the present disclosure, the user's fourth operation may be a triggering operation on a storage control. Exemplarily, the triggering operation on the storage control may be clicking a storage control trigger button on the function list interface. For example, when the user Figure 3 opens the structure storage sliding switch button on the shown function list interface, the intelligent assistant can convert the current document (ppt, word, pdf, jpg, etc.) into a standard format such as Markdown or Json for saving. Among them, the text content can be saved separately as a Markdown text or a Json file, and the charts in the document can be saved separately in one or more folders, and the paths are written in the corresponding positions in the Markdown text or the Json file. Through structured storage, the intelligent assistant can embed the content of the recommended reference resources into the document or the document picture (such as as a comment or directly embed the relevant content of the document during standardization).
[0154] In the embodiments of the present disclosure, by using technologies such as layout analysis technology, natural language processing technology, data storage and retrieval, etc., a document processing assistant with intelligent recommendation and personalized recommendation functions is realized, which greatly improves the user experience and information acquisition efficiency.
[0155] Embodiments of the present disclosure can implement intelligent reference material recommendations related to document content for document processing scenarios such as business contracts, legal documents, academic papers, project reports, etc. For example, for business contract or legal document review scenarios, using the resource recommendation method of the embodiments of the present disclosure, the clauses that need to be reviewed key points can be automatically analyzed according to the contract structure, or legal cases can be recommended. For academic paper reading / writing scenarios, using the resource recommendation method of the embodiments of the present disclosure, relevant literature, materials or editing suggestions can be automatically recommended according to the chapter structure of the paper. For project report writing scenarios, using the resource recommendation method of the embodiments of the present disclosure, relevant reference materials can be recommended according to the structure of the project report, and users can conduct comparative analysis with other relevant projects based on the relevant reference materials. However, the application scenarios of the present disclosure are not limited thereto. The resource recommendation method of the embodiments of the present disclosure can intelligently recommend the required reference resources to users when they process documents, facilitating user operations and greatly improving the efficiency of users in processing documents.
[0156] As Figure 8 shown, embodiments of the present disclosure also provide a resource recommendation device, including a data acquisition module 810, a layout analysis module 820, and a recommendation module 830, where:
[0157] The data acquisition module 810 is configured to obtain a document picture or receive a document input by the user in response to a first operation of the user;
[0158] The layout analysis module 820 is configured to divide the document picture or document into one or more modules and detect the content of the one or more modules;
[0159] The recommendation module 830 is configured to extract semantic vectors of at least one module according to the content of the detected module, and select one or more of the reference resources from the semantic vectors of a plurality of pre-stored reference resources as the recommended content corresponding to the module.
[0160] In some exemplary embodiments, the first operation is a trigger operation on a screenshot control or a trigger operation on an import control, and the trigger operation on the screenshot control includes at least one of the following:
[0161] After opening the document, click the screenshot control trigger button on the function list interface;
[0162] Open the document;
[0163] After the position of the scroll bar on the document operation interface changes and then stops and the stop duration reaches a preset stop duration threshold;
[0164] The trigger operation on the import control includes at least one of the following:
[0165] Click the import control trigger button on the function list interface;
[0166] Drag the document into the first icon, where the function list interface is the interface presented after operating on the first icon.
[0167] In some exemplary embodiments, the recommendation module 830 is further configured to display the recommended content of the corresponding module in a pop-up window or a collapsible panel.
[0168] In some exemplary embodiments, the resource recommendation device further includes a communication module, and the communication module is configured to, in response to a second operation of the user, obtain the upper and lower boundaries of the visible area of the document operation interface, where the second operation includes: opening the document, and / or, after the position of the scroll bar of the document operation interface changes and stops and the stop duration reaches a preset stop duration threshold;
[0169] The recommendation module 830 is further configured to determine the position of the detected module between the upper and lower boundaries of the visible area of the document operation interface, and determine the display position of the pop-up window according to the position of the module.
[0170] In some exemplary embodiments, the recommendation module 830 is further configured to perform at least one of the following:
[0171] In response to a drag operation of the user on the pop-up window, move the position of the pop-up window;
[0172] In response to detecting that the size of the document content covered by the pop-up window exceeds a preset area threshold, automatically move the position of the pop-up window according to the document content;
[0173] In response to a window size adjustment operation of the user on the pop-up window, adjust the size of the pop-up window;
[0174] In response to detecting an interaction operation of the user on the pop-up window, automatically adjust the size of the pop-up window according to the size of the recommended content;
[0175] In response to a transparency parameter setting operation of the user on the pop-up window, adjust the transparency of the pop-up window;
[0176] In response to detecting that the size of the document content covered by the pop-up window exceeds a preset area threshold, automatically adjust the transparency of the pop-up window.
[0177] In some exemplary embodiments, the document includes at least one of the following: a pure picture document, a pure text document, a Word document, a PDF document, a PPT document, a CAJ document.
[0178] In some exemplary embodiments, the document picture is determined according to the upper and lower boundaries of the visible area of the current document operation software, or according to the user's screenshot interface.
[0179] In some exemplary embodiments, one or more modules include at least one of the following types: pictures, tables, formulas, text, explanatory text, titles, references.
[0180] In some exemplary embodiments, the layout analysis module 820 divides the document picture or document into one or more modules, including:
[0181] Input the document picture or document into a layout detection model to obtain the division result of the document picture or document. Wherein, when the number of pages of the document picture or the number of pages of the document is less than or equal to a preset page number threshold, the layout detection model is a local model; when the number of pages of the document picture or the number of pages of the document is greater than the preset page number threshold, the layout detection model is a cloud model.
[0182] In some exemplary embodiments, the layout analysis module 820 detects the content of the one or more modules, including:
[0183] Select at least one of the modules;
[0184] Determine a corresponding content recognition model according to the type of the selected module;
[0185] Input the selected module into the corresponding content recognition model to obtain the content of the module.
[0186] In some exemplary embodiments, the recommendation module 830 extracts the semantic vector of at least one module according to the detected content of the module, including:
[0187] Extract the first information of the module according to the detected content of the module;
[0188] Input the extracted first information into a first pre-trained model to obtain the semantic vector of the module.
[0189] In some exemplary embodiments, the recommendation module 830 selects one or more of the reference resources from the semantic vectors of multiple pre-stored reference resources as the recommended content for the corresponding module according to the extracted semantic vector, including any one of the following:
[0190] Calculate the similarity between the extracted semantic vector and the semantic vectors of multiple pre-stored reference resources, and select one or more of the reference resources as the recommended content for the corresponding module according to the calculation result of the similarity between the extracted semantic vector and the semantic vectors of multiple reference resources;
[0191] Input the extracted semantic vectors into the trained recommendation model to obtain the reference resources output by the recommendation model, and use the reference resources output by the recommendation model as the recommended content for the corresponding module, where the recommendation model is pre-trained using the semantic vectors of the multiple pre-stored reference resources;
[0192] Calculate the distances between the extracted semantic vectors and multiple class center vectors, and select one or more reference resources as the recommended content for the corresponding module among the classes where the calculated distances are lower than the preset distance threshold, where the multiple class center vectors are obtained by performing clustering analysis on the semantic vectors of the multiple pre-stored reference resources.
[0193] In some exemplary embodiments, the recommendation module 830 selects one or more of the reference resources as the recommended content for the corresponding module according to the similarity calculation results between the extracted semantic vectors and the semantic vectors of the multiple reference resources, including:
[0194] Select the reference resources with similarity calculation results higher than the preset similarity threshold as the recommended content for the corresponding module.
[0195] In some exemplary embodiments, the recommendation module 830 selects one or more of the reference resources as the recommended content for the corresponding module according to the similarity calculation results between the extracted semantic vectors and the semantic vectors of the multiple reference resources, including:
[0196] Obtain the user's interest field and the domain feature vectors of multiple reference resources;
[0197] Calculate the domain weight of each reference resource according to the user's interest field and the domain feature vectors of the multiple reference resources;
[0198] Calculate the weighted similarity calculation result of each reference resource according to the similarity calculation result and the domain weight of each reference resource;
[0199] Select the reference resources with weighted similarity calculation results higher than the preset similarity threshold as the recommended content for the corresponding module.
[0200] In some exemplary embodiments, the recommendation module 830 calculates the domain weight of each reference resource according to the user's interest field and the domain feature vectors of the multiple reference resources, including any one of the following:
[0201] Calculate the vector inner product between the domain feature vector of each reference resource and the user interest weight, and use the calculated vector inner product as the domain weight of each reference resource;
[0202] Calculate the similarity between the domain feature vector of each of the reference resources and the user interest weights, and use the calculation result of the similarity between the domain feature vector of each of the reference resources and the user interest weights as the domain weight of each of the reference resources.
[0203] An embodiment of the present disclosure further provides a resource recommendation device, including a memory; and a processor connected to the memory, the memory is used to store instructions, and the processor is configured to execute the steps of the resource recommendation method as described in any embodiment of the present disclosure based on the instructions stored in the memory.
[0204] As Figure 9 shown, in one example, the resource recommendation device may include: a processor 910, a memory 920, a bus system 930, and a transceiver 940. Among them, the processor 910, the memory 920, and the transceiver 940 are connected through the bus system 930. The memory 920 is used to store instructions, and the processor 910 is used to execute the instructions stored in the memory 920 to control the transceiver 940 to transmit and receive signals. Specifically, the transceiver 940 can, under the control of the processor 910, in response to a first operation of the user, acquire a document picture or receive a document input by the user. The processor 910 divides the document picture or the document into one or more modules, and detects the content of the one or more modules; extracts the semantic vector of at least one module according to the detected content of the module, and selects one or more of the reference resources from the semantic vectors of a plurality of pre-stored reference resources as the recommended content corresponding to the module.
[0205] It should be understood that the processor 910 may be a central processing unit (CPU), and the processor 910 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0206] The memory 920 may include a read-only memory and a random access memory, and provide instructions and data to the processor 910. A part of the memory 920 may also include a non-volatile random access memory. For example, the memory 920 may also store information about the device type.
[0207] The bus system 930 may include, in addition to the data bus, a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, in Figure 9 all kinds of buses are labeled as the bus system 930.
[0208] In the implementation process, the processing performed by the processing device can be completed by the integrated logic circuit of the hardware in the processor 910 or the instructions in the form of software. That is, the method steps of the embodiments of the present disclosure can be embodied as being executed and completed by the hardware processor, or by a combination of the hardware and software modules in the processor. The software module can be located in a storage medium such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 920, and the processor 910 reads the information in the memory 920 and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0209] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the resource recommendation method as described in any embodiment of the present disclosure. By executing the executable instructions, the driven resource recommendation method is basically the same as the resource recommendation method provided in the above embodiments of the present disclosure, and will not be elaborated here.
[0210] In some possible implementation manners, each aspect of the resource recommendation method provided by the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in the resource recommendation method according to various exemplary embodiments of the present disclosure described in this specification. For example, the computer device can execute the resource recommendation method recorded in the embodiments of the present disclosure.
[0211] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0212] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof. In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, one physical component may have multiple functions, or one function or step may be executed by several physical components in cooperation. Some or all of the components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0213] It should be noted that the above embodiments or implementations are merely exemplary and not restrictive. Therefore, the present disclosure is not limited to the content specifically shown and described herein. Various modifications, substitutions, or omissions can be made to the form and details of the implementation without departing from the scope of the present disclosure.
Claims
1. A resource recommendation method, characterized in that, Including: In response to a first operation of the user, obtaining a document picture or receiving a document input by the user; Dividing the document picture or document into one or more modules, and detecting the content of the one or more modules; Extracting semantic vectors of at least one module according to the content of the detected modules, and selecting one or more of the reference resources from the semantic vectors of multiple pre-stored reference resources as the recommended content corresponding to the modules.
2. The method according to claim 1, wherein The first operation is a triggering operation on a screenshot control or a triggering operation on an import control. The triggering operation on the screenshot control includes at least one of the following: After opening the document, clicking the triggering button of the screenshot control on the function list interface; Opening the document; After the position of the scroll bar on the document operation interface changes and then stops, and the stop duration reaches a preset stop duration threshold; The triggering operation on the import control includes at least one of the following: Clicking the triggering button of the import control on the function list interface; Dragging the document into the first icon, where the function list interface is the interface presented after operating on the first icon.
3. The method according to claim 1, wherein The method further includes: In response to a second operation of the user, obtaining the upper and lower boundaries of the visible area of the document operation interface. The second operation includes: opening the document, and / or, after the position of the scroll bar on the document operation interface changes and then stops, and the stop duration reaches a preset stop duration threshold; Determining the position of the detected module between the upper and lower boundaries of the visible area of the document operation interface, and determining the display position of the recommended content according to the position of the module.
4. The method according to claim 1, wherein The method further includes: displaying the recommended content corresponding to the module in a pop-up window or a collapsible panel.
5. The method according to claim 4, wherein When displaying the recommended content corresponding to the module in the pop-up window, the method further includes at least one of the following: In response to a dragging operation of the user on the pop-up window, moving the position of the pop-up window; In response to detecting that the size of the document content covered by the pop-up window exceeds a preset area threshold, automatically moving the position of the pop-up window according to the document content; In response to a window size adjustment operation of the user on the pop-up window, adjusting the size of the pop-up window; In response to detecting an interaction operation of the user on the pop-up window, automatically adjusting the size of the pop-up window according to the size of the recommended content; In response to a transparency parameter setting operation of the user on the pop-up window, adjusting the transparency of the pop-up window; In response to detecting that the size of the document content covered by the pop-up window exceeds a preset area threshold, automatically adjusting the transparency of the pop-up window.
6. The method according to claim 1, wherein The one or more modules include at least one of the following types: picture, table, formula, text, description text, title, reference.
7. The method according to claim 1, characterized in that, The dividing the document picture or document into one or more modules includes: Inputting the document picture or document into a layout detection model to obtain the division result of the document picture or document. When the number of pages of the document picture or the number of pages of the document is less than or equal to a preset page threshold, the layout detection model is a local model; when the number of pages of the document picture or the number of pages of the document is greater than the preset page threshold, the layout detection model is a cloud model.
8. The method according to claim 1, wherein The detecting the content of the one or more modules includes: Select at least one of the said modules; Determine the corresponding content recognition model according to the type of the selected module; Input the selected module into the corresponding content recognition model to obtain the content of the module.
9. The method according to claim 1, wherein The extracting at least one semantic vector of the module according to the detected content of the module includes: Extract the first information of the module according to the detected content of the module; Input the extracted first information into the first pre-trained model to obtain the semantic vector of the module.
10. The method according to claim 1, characterized in that, The selecting one or more of the reference resources from the semantic vectors of the multiple pre-stored reference resources as the recommended content for the corresponding module according to the extracted semantic vector includes any one of the following: Calculate the similarity between the extracted semantic vector and the semantic vectors of the multiple pre-stored reference resources, and select one or more of the reference resources as the recommended content for the corresponding module according to the calculation result of the similarity between the extracted semantic vector and the semantic vectors of the multiple reference resources; Input the extracted semantic vector into the trained recommendation model to obtain the reference resources output by the recommendation model, and use the reference resources output by the recommendation model as the recommended content for the corresponding module, where the recommendation model is pre-trained using the semantic vectors of the multiple pre-stored reference resources; Calculate the distance between the extracted semantic vector and multiple class center vectors, and select one or more reference resources as the recommended content for the corresponding module from the classes in which the calculated distance is lower than the preset distance threshold, where the multiple class center vectors are obtained by performing clustering analysis on the semantic vectors of the multiple pre-stored reference resources.
11. The method according to claim 10, characterized in that The selecting one or more of the reference resources from the semantic vectors of the multiple reference resources as the recommended content for the corresponding module according to the calculation result of the similarity between the extracted semantic vector and the semantic vectors of the multiple reference resources includes: Select the reference resources with the similarity calculation result higher than the preset similarity threshold as the recommended content for the corresponding module.
12. The method according to claim 10, characterized in that, The selecting one or more of the reference resources from the semantic vectors of the multiple reference resources as the recommended content for the corresponding module according to the calculation result of the similarity between the extracted semantic vector and the semantic vectors of the multiple reference resources includes: Obtain the user's interest field and the field feature vectors of the multiple reference resources; Calculate the field weight of each reference resource according to the user's interest field and the field feature vectors of the multiple reference resources; Calculate the weighted similarity calculation result of each reference resource according to the similarity calculation result and the field weight of each reference resource; Select the reference resources with the weighted similarity calculation result higher than the preset similarity threshold as the recommended content for the corresponding module.
13. The method according to claim 12, wherein The calculating the field weight of each reference resource according to the user's interest field and the field feature vectors of the multiple reference resources includes any one of the following: Calculate the vector inner product between the field feature vector of each reference resource and the user interest weight, and use the calculated vector inner product as the field weight of each reference resource; Calculate the similarity between the domain feature vector of each of the reference resources and the user interest weights, and use the calculation result of the similarity between the domain feature vector of each of the reference resources and the user interest weights as the domain weight of each of the reference resources.
14. A resource recommendation device, characterized in that, Comprising a memory; and a processor connected to the memory, the memory being configured to store instructions, the processor being configured to execute the steps of the resource recommendation method according to any one of claims 1 to 13 based on the instructions stored in the memory.
15. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, the resource recommendation method according to any one of claims 1 to 13 is implemented.
16. A computer program product, characterized in that, Comprising instructions, when the computer program product is executed by a computer, the instructions execute the resource recommendation method according to any one of claims 1 to 13.
17. A resource recommendation device, characterized in that, Comprising a data acquisition module, a layout analysis module and a recommendation module, wherein: The data acquisition module is configured to acquire a document picture or receive a document input by a user in response to a first operation of the user; The layout analysis module is configured to divide the document picture or the document into one or more modules and detect the content of the one or more modules; The recommendation module is configured to extract the semantic vector of at least one module according to the content of the detected module, and select one or more of the reference resources from the semantic vectors of a plurality of pre-stored reference resources as the recommended content corresponding to the module.
Citation Information
Cited By
Non-modal intelligent assistant interactive interface system for industrial design software
CN122064417A
A non-modal intelligent assistant interface system for industrial design software
CN122064417B