Content extraction methods and their application in developing creative responses
By identifying and classifying content objects, and combining image-to-text technology with AI natural language processing, the problem of content extraction and response in R&D creative scenarios has been solved, providing flexible and highly referential response results and stimulating researchers' creativity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies lack effective content extraction and response methods in creative research scenarios, especially in scenarios with strong subjectivity or extensibility, making it difficult to provide referential and logically sound response results.
By acquiring content objects, identifying and classifying content elements, combining image-to-text recognition, using a trained detection neural network to generate image-to-text data, extracting keywords, and combining keywords and derived vocabulary to form question statements, which are then input into AI natural language processing tools for response.
It enables efficient and flexible content extraction and response in R&D and creative scenarios, provides highly referential interactive results, and inspires researchers' secondary inspiration and technical ideas.
Smart Images

Figure CN116737905B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of content extraction technology and information response technology, and in particular to content extraction methods and their application in the development of creative responses. Background Technology
[0002] With the increasing adoption of AI technology, the application scenarios for extracting text content using AI-trained language processing tools are becoming more and more widespread. Especially in some template-based scenarios, the use of trained language processing tools (such as ChatGPT and Wenxin Yiyan) for content extraction can save users a lot of repetitive work time. However, for some more subjective or extensible scenarios, there is little literature on the use of language processing tools for assistance, such as information collection or expansion of R&D ideas. Therefore, if a solution can be provided for R&D idea scenarios that can output reference-oriented and logically sound response results, it will provide positive assistance in stimulating and expanding R&D ideas. Summary of the Invention
[0003] In view of this, the purpose of this invention is to propose a content extraction method that is reliable in implementation, flexible in operation, and provides highly referential response results, and its application in the research and development of creative responses.
[0004] To achieve the above-mentioned technical objectives, the technical solution adopted by this invention is as follows:
[0005] A content extraction method, comprising:
[0006] S01. Obtain the first content object, identify its content elements, and generate the identification result;
[0007] S02. Based on the recognition results, the first content object is classified into image data and text data.
[0008] S03. Import the image portion data into the trained detection neural network for content recognition, and the detection neural network generates at least one image-to-text data corresponding to the image portion data.
[0009] S04. Extract keywords from the text data and image-to-text data according to preset requirements to complete the content extraction results.
[0010] As a possible implementation, further, in this solution S01, when the first content object is a handwritten manuscript, before recognizing its content elements, it is also scanned to form an electronic image file. Then, the formed electronic image file is denoised and its contrast is adjusted according to preset conditions. Finally, its content elements are recognized to generate a recognition result.
[0011] As a preferred implementation method, preferably, in this solution S01, the recognition result includes the contour data of each content element in the first content object and its corresponding image or text marking information.
[0012] As a preferred implementation method, in S03 of this solution, the training method for the trained detection neural network is preferably as follows:
[0013] S031. Obtain a preset amount of image data, then mark the content elements in the image data with information, then associate the information marks with the image data to form training data, and finally aggregate the preset amount of training data to form a training dataset.
[0014] S032. Extract different amounts of training data from the training dataset as training and validation groups, and then import them into the neural network for training until the model converges to obtain the detection neural network.
[0015] Based on the above, the present invention also provides a method for responding to creative information, which includes the content extraction method described above, and further includes:
[0016] A01. Respond to the R&D creative information input by the user and set it as the first content object;
[0017] A02, execute S01~S04;
[0018] A03. Obtain the keywords corresponding to the content extraction results, and then perform derivative processing on them to obtain derivative words related to the keywords;
[0019] A04. Based on keywords and derived words, combine one or more keywords or one or more keywords with one or more derived words to form a word group;
[0020] A05. Using vocabulary groups as core words, insert them into a preset response request template to form a question statement;
[0021] A06. Import the question into an AI-trained natural language processing tool to obtain the response results;
[0022] A07. Associate the response results with the question and the corresponding vocabulary group to form R&D creative information response data and complete the response interaction.
[0023] As a preferred implementation method, preferably, in solution A03, the method for keyword derivation processing includes:
[0024] A031. Using keywords as initial vocabulary, obtain their synonyms, related words, and related terms, and then use them as first-level derived vocabulary.
[0025] A032. Derive from the first-level derived words, obtain their synonyms, related words, and hyponyms, and then use them as second-level derived words.
[0026] As a preferred implementation method, the keywords and derived words in this solution are associated in the form of a mind map, wherein the keywords are the main items and their corresponding derived words are the branch items of the main items, and the derived words are also labeled with the derivative level during the derivative processing.
[0027] As a preferred implementation method, solution A06 further includes: using the vocabulary group as the core word, importing it into a search engine for searching, obtaining the top N pieces of information that meet the preset relevance requirements, and then extracting their links;
[0028] A07 also includes: associating the links corresponding to the vocabulary groups with them.
[0029] Based on the above, the present invention also provides a creative information response system, which includes:
[0030] The content acquisition unit is used to respond to the R&D creative information input by the user and set it as the first content object;
[0031] The content recognition unit is used to acquire the first content object, recognize its content elements, and generate recognition results.
[0032] The content classification unit is used to classify the first content object into image data and text data based on the recognition results.
[0033] A detection neural network unit is used to perform content recognition on partial image data and generate at least one image-to-text data corresponding to the partial image data.
[0034] The content extraction unit is used to extract keywords from text data and image-to-text data according to preset requirements, and complete the content extraction results.
[0035] The content derivation unit is used to obtain the keywords corresponding to the content extraction results, and then perform derivation processing on them to obtain derivative words related to the keywords;
[0036] Content combination unit, used to combine one or more keywords or one or more keywords with one or more derivative words to form a word group based on keywords and derivative words;
[0037] The response processing unit is used to take a vocabulary group as the core word, put it into a preset response request template to form a question statement; and also import the question statement into an AI-trained natural language processing tool to obtain the response result.
[0038] The data association unit is used to associate the response results with the question statement and the corresponding vocabulary group to form R&D creative information response data and complete the response interaction.
[0039] Based on the above, the present invention also provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the content extraction method or the R&D creative information response method described above.
[0040] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: The ingenuity of this solution is to extract the R&D creative information into a first content object, and then combine it with content element recognition to extract the first content object into images and text. Then, by combining image-to-text recognition, the image information in the R&D creative information is also converted into text. Finally, by integrating the recognized content into a question statement and inputting it into an AI-trained natural language processing tool for response, the natural language processing tool can better respond to the relevant content. Users can combine the response results to expand and stimulate secondary inspiration on the input R&D creative information, providing researchers with better technical ideas for R&D creativity. This solution is not only simple to implement but also has good reference value. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a simplified implementation flowchart of the content extraction method in this solution;
[0043] Figure 2 This is a simplified implementation flowchart illustrating the application of the content extraction method of this solution in the development of creative responses;
[0044] Figure 3 This is a simplified diagram illustrating the content extraction and vocabulary derivation process for creative research content in this solution;
[0045] Figure 4 This is a simplified diagram showing the connection of the unit modules of this system. Detailed Implementation
[0046] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] like Figure 1 As shown, this embodiment of the solution provides a content extraction method, which includes:
[0048] S01. Obtain the first content object, identify its content elements, and generate the identification result;
[0049] S02. Based on the recognition results, the first content object is classified into image data and text data.
[0050] S03. Import the image portion data into the trained detection neural network for content recognition, and the detection neural network generates at least one image-to-text data corresponding to the image portion data.
[0051] S04. Extract keywords from the text data and image-to-text data according to preset requirements to complete the content extraction results.
[0052] To improve the applicability of the solution, the original information of the first content object in this solution can be a printed manuscript (such as a printed copy of an electronic document), but it is not limited to this. It can also be a handwritten manuscript. As a possible implementation, in S01 of this solution, when the first content object is a handwritten manuscript, it is scanned before its content element recognition to form an electronic image file. Then, the formed electronic image file is denoised and its contrast is adjusted according to preset conditions. Finally, its content element is recognized to generate a recognition result.
[0053] Denoising involves removing interference spots or small dots / color blocks that have no substantial content from the scanned file.
[0054] Regarding content recognition, as a preferred implementation method, in S01 of this solution, the recognition result includes the contour data of each content element in the first content object and its corresponding image or text marking information; wherein, in terms of content element recognition, this solution can locate the element contour through a trained localization neural network, and then extract and recognize it through a detection neural network. Since the neural network applied to image and text localization or detection can be directly implemented using currently published literature, it will not be described in detail.
[0055] For the recognition of partial image data, as a preferred implementation method, in this scheme S03, the training method of the trained detection neural network is as follows:
[0056] S031. Obtain a preset amount of image data, then mark the content elements in the image data with information, then associate the information marks with the image data to form training data, and finally aggregate the preset amount of training data to form a training dataset.
[0057] S032. Extract different amounts of training data from the training dataset as training and validation groups, and then import them into the neural network for training until the model converges to obtain the detection neural network.
[0058] Combination Figure 2 As shown, based on the above, this implementation scheme also provides a method for responding to creative information, which includes the content extraction method described above, and further includes:
[0059] A01. Respond to the R&D creative information input by the user and set it as the first content object;
[0060] A02, execute S01~S04;
[0061] A03. Obtain the keywords corresponding to the content extraction results, and then perform derivative processing on them to obtain derivative words related to the keywords;
[0062] A04. Based on keywords and derived words, combine one or more keywords or one or more keywords with one or more derived words to form a word group;
[0063] A05. Using vocabulary groups as core words, insert them into a preset response request template to form a question statement;
[0064] A06. Import the question into an AI-trained natural language processing tool to obtain the response results;
[0065] A07. Associate the response results with the question and the corresponding vocabulary group to form R&D creative information response data and complete the response interaction.
[0066] As a preferred implementation method, preferably, in solution A03, the method for keyword derivation processing includes:
[0067] A031. Using keywords as initial vocabulary, obtain their synonyms, related words, and related terms, and then use them as first-level derived vocabulary.
[0068] A032. Derive from the first-level derived words, obtain their synonyms, related words, and hyponyms, and then use them as second-level derived words.
[0069] Combination Figure 3 As shown, as a preferred implementation method, the keywords and derived words in this solution are associated in the form of a mind map. The keywords are the main items, and their corresponding derived words are the branch items of the main items. In addition, the derived words are also labeled with the derivation level during the derivation process.
[0070] In addition, this solution A06 also includes: using vocabulary groups as core words, importing them into a search engine for searching, obtaining the top N pieces of information that meet the preset relevance requirements, and then extracting their links;
[0071] A07 also includes: associating the links corresponding to the vocabulary groups with them.
[0072] Combination Figure 4 As shown, based on the above, this implementation scheme also provides a research and development creative information response system, which includes:
[0073] The content acquisition unit is used to respond to the R&D creative information input by the user and set it as the first content object;
[0074] The content recognition unit is used to acquire the first content object, recognize its content elements, and generate recognition results.
[0075] The content classification unit is used to classify the first content object into image data and text data based on the recognition results.
[0076] A detection neural network unit is used to perform content recognition on partial image data and generate at least one image-to-text data corresponding to the partial image data.
[0077] The content extraction unit is used to extract keywords from text data and image-to-text data according to preset requirements, and complete the content extraction results.
[0078] The content derivation unit is used to obtain the keywords corresponding to the content extraction results, and then perform derivation processing on them to obtain derivative words related to the keywords;
[0079] Content combination unit, used to combine one or more keywords or one or more keywords with one or more derivative words to form a word group based on keywords and derivative words;
[0080] The response processing unit is used to take a vocabulary group as the core word, put it into a preset response request template to form a question statement; and also import the question statement into an AI-trained natural language processing tool to obtain the response result.
[0081] The data association unit is used to associate the response results with the question statement and the corresponding vocabulary group to form R&D creative information response data and complete the response interaction.
[0082] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0083] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0084] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A research and development creative information response method characterized by comprising: It comprises: A01, in response to the user inputted research and development creative information content, set it as the first content object; A02, execute content extraction according to preset conditions, generate content extraction result; A03, obtain the keywords corresponding to the content extraction result, and then derive the keywords to obtain the derivative vocabulary related to the keywords; A04, according to the keywords and the derivative vocabulary, combine more than one keyword or more than one keyword with more than one derivative vocabulary to form a vocabulary group; A05, taking the vocabulary group as the core word, put it into the preset answer request template to form a question sentence; A06, import the question sentence into the natural language processing tool trained by AI to obtain the answer result; A07, associate the answer result with the question sentence and the corresponding vocabulary group to form the research and development creative information answer data, and complete the answer interaction; Wherein, the content extraction comprises: S01, obtain the first content object, and perform content element recognition to generate an identification result, wherein when the first content object is a handwritten manuscript, it is scanned to form an electronic image file before content element recognition, and then the electronic image file is denoised and contrast adjusted according to preset conditions, and finally the content element recognition is performed to generate the identification result; The identification result includes the contour data of each content element in the first content object and the corresponding image or text mark information; S02, according to the identification result, the first content object is classified into image part data and text part data; S03, import the image part data into the trained detection neural network for content recognition, and generate at least one image-to-text data corresponding to the image part data by the detection neural network; S04, extract keywords from the text part data and image-to-text data according to preset requirements to complete the content extraction result.
2. The R&D idea information response method according to claim 1, wherein In S03, the training method of the trained detection neural network is: S031, obtain a preset amount of image data, then mark the information of the content elements in the image data, then associate the information mark with the image data to form training data, and finally collect a preset amount of training data to form a training data set; S032, respectively extract a preset amount of different training data from the training data set as a training group and a validation group, then import them into the neural network for training until the model converges, and obtain the detection neural network.
3. The R&D idea information response method according to claim 1, wherein In A03, the method for deriving the keywords comprises: A031, taking the keyword as the initial vocabulary, obtaining its near-synonyms, synonyms and superordinate and subordinate concept words, and then taking them as first-level derivative vocabulary; A032, derive the first-level derivative vocabulary to obtain its near-synonyms, synonyms and superordinate and subordinate concept words, and then take them as second-level derivative vocabulary.
4. The R&D idea information response method of claim 1, wherein, The keywords and derivative vocabulary are associated in the form of mind map, wherein the keyword is the main project, the derivative vocabulary corresponding to the keyword is the branch project of the main project, and the derivative vocabulary is also labeled with derivative level during the derivation process.
5. The R&D idea information response method according to claim 1, wherein A06 further comprises: taking the vocabulary group as a core word, importing it into a search engine for searching, and obtaining the top N information with relevance compound preset requirements, and then extracting the links thereof; A07 further comprises: associating the links corresponding to the vocabulary group with the vocabulary group.
6. A development idea information response system which applies the development idea information response method according to any one of claims 1 to 5, characterized by It comprises: A content acquisition unit for responding to the research and development creative information content input by the user and setting it as a first content object; A content recognition unit for obtaining the first content object, performing content element recognition thereon, and generating a recognition result; A content classification unit for classifying the first content object into image part data and text part data according to the recognition result; A detection neural network unit for performing content recognition on the image part data and generating at least one image-to-text data corresponding to the image part data; A content extraction unit for extracting keywords from the text part data and the image-to-text data according to preset requirements and completing a content extraction result; A content derivation unit for obtaining keywords corresponding to the content extraction result and then performing derivation processing thereon to obtain derivative words related to the keywords; A content combination unit for combining one or more keywords or one or more keywords and one or more derivative words to form a vocabulary group according to the keywords and the derivative words; A response processing unit for taking the vocabulary group as a core word, placing it in a preset response request template to form a question sentence, and further importing the question sentence into an AI-trained natural language processing tool to obtain a response result; A data association unit for associating the response result with the question sentence and the corresponding vocabulary group to form research and development creative information response data and complete response interaction.
7. A computer readable storage medium characterized by: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, which is loaded and executed by the processor to realize the research and development creative information response method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Visual question and answer method and device and electronic equipment
CN114186039A
Intention recognition model training method and dialogue method and device
CN115345177A