Information processing method and device
By identifying the target material content and fragment objects in the display interface, generating prompt words and inputting them into the artificial intelligence model, the problem of the artificial intelligence model being unable to adapt to the user's personalized tasks is solved, thus improving the accuracy and reliability of information reasoning.
Patent Information
- Application Number
- CN202511494456.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-02-17
AI Technical Summary
Existing artificial intelligence models cannot adapt to users' specific personalized task needs, resulting in low accuracy of inference results that deviate from the user's actual intentions.
By responding to the user's material selection operation, the target material content is determined in the display interface, the user input information is obtained, the target segment object is determined from the material content, and prompt words are generated and input into the artificial intelligence model for information processing.
It enables flexible access to different user-selected content for cross-modal information integration, accurately matching user intent, improving the accuracy and reliability of information reasoning, and meeting personalized task requirements.
Smart Images

Figure CN121543710A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an information processing method and apparatus. Background Technology
[0002] With the rapid development of computer technology, various artificial intelligence models have begun to be widely used in academia and industry. Among them, artificial intelligence models, trained on deep neural network architectures and massive unlabeled corpora, can model the complex patterns of human language. Specifically, based on multimodal semantic association mapping, they can complete multiple tasks such as cross-language translation, information summarization, problem reasoning and solving, and code symbol generation, providing efficient technical support for application scenarios such as information retrieval, content creation, and human-computer collaboration.
[0003] Related technologies employ a general knowledge base and apply artificial intelligence models to perform information reasoning based on this knowledge base to meet diverse task requirements. However, this approach, where the AI model performs reasoning and analysis based on general data, cannot adapt to the specific personalized task needs of users, resulting in deviations from the user's actual intent and low accuracy of the reasoning results. Summary of the Invention
[0004] This application provides an information processing method and apparatus that solves the problem in related technologies that, when performing information reasoning, cannot adapt to the specific personalized task needs of users, deviate from the actual intentions of users, and have low accuracy of reasoning results. It can flexibly call different material contents selected by users to integrate cross-modal information, effectively match the actual intentions of users for information reasoning, meet the specific personalized task needs of users, and improve the accuracy and reliability of reasoning results.
[0005] In a first aspect, embodiments of this application provide an information processing method, the method comprising: In response to the material selection operation, at least one piece of target material content is selected in the currently displayed interface; Obtain user input information, and determine at least one target fragment object from the at least one segment of target material content based on the user input information; Based on the user input information and the at least one target fragment object, the prompt word information is determined, and the prompt word information is input into the set artificial intelligence model to obtain the information processing result.
[0006] Secondly, embodiments of this application also provide an information processing apparatus, including: The material content determination module is configured to determine at least one piece of target material content in the currently displayed interface in response to the material selection operation; The target segment determination module is configured to acquire user input information and determine at least one target segment object from the at least one segment of target material content based on the user input information; The processing result determination module is configured to determine prompt word information based on the user input information and the at least one target fragment object, and input the prompt word information into a set artificial intelligence model to obtain the information processing result.
[0007] Thirdly, embodiments of this application also provide an information processing device, which includes: One or more processors; Storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the information processing method described in the embodiments of this application.
[0008] Fourthly, embodiments of this application also provide a non-volatile storage medium for storing computer-executable instructions, which, when executed by a computer processor, are configured to perform the information processing method described in embodiments of this application.
[0009] In this embodiment, in response to a material selection operation, at least one segment of target material content is determined in the current display interface. This allows users to quickly select the required material content through intuitive interface operations and accurately locate the material range. Based on user input information, at least one target fragment object is determined from the at least one segment of target material content, accurately filtering material fragments that match the user's intent. Prompt word information is determined based on user input information and at least one target fragment object, and this prompt word information is input into a set artificial intelligence model to obtain information processing results. This effectively integrates fragmented information and improves the accuracy of the model's output. The above solution can flexibly call upon different material contents selected by the user for cross-modal information integration, effectively match the user's actual intent for information reasoning, meet the user's specific personalized task needs, and improve the accuracy and reliability of the reasoning results. Attached Figure Description
[0010] Figure 1 A flowchart illustrating an information processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating how to determine target material content by selecting words using a word selection box, as provided in an embodiment of this application. Figure 3 This is a schematic diagram illustrating how to determine the content of target material by selecting a frame in a screenshot, as provided in an embodiment of this application. Figure 4 A flowchart illustrating a process for determining target material content, provided in an embodiment of this application; Figure 5 A schematic diagram illustrating how a fragment control is triggered in a dialog box area, as provided in an embodiment of this application; Figure 6 A flowchart of an information processing method including a process for determining a target fragment object is provided in an embodiment of this application; Figure 7 A flowchart illustrating a process for determining candidate fragment objects is provided in an embodiment of this application; Figure 8 A flowchart illustrating a process for retrieving and determining a target fragment object, provided as an embodiment of this application; Figure 9 A flowchart of an information processing method including a process of determining prompt word information is provided for an embodiment of this application; Figure 10 A structural block diagram of an information processing device provided in an embodiment of this application; Figure 11 This is a schematic diagram of the structure of an information processing device provided in an embodiment of this application. Detailed Implementation
[0011] The embodiments of this application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of the embodiments of this application and are not intended to limit the scope of the embodiments. Furthermore, it should be noted that, for ease of description, only the parts relevant to the embodiments of this application are shown in the accompanying drawings, not the entire structure.
[0012] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0013] The information processing method provided in this application is used to integrate and recall information by combining different material content selected by the user, and to accurately match the user's intent for information reasoning, which can meet the user's different task needs. For example, writing, data analysis, data summarization, and knowledge question answering. The aforementioned application scenarios are merely exemplary and illustrative. In practical applications, this information processing method can also be used in other information processing scenarios, and this application does not limit this. This application aims to provide an information processing method and apparatus to solve the problem in related technologies that, when performing information reasoning, cannot adapt to the user's specific personalized task needs, deviate from the user's actual intent, and have low accuracy in reasoning results.
[0014] The information processing method provided in this application embodiment can be executed by a computer device. The computer device refers to any electronic device with data computing, processing and storage capabilities, such as mobile phones, PCs (Personal Computers), tablet computers and other terminal devices, or servers and other devices. This application embodiment does not limit the scope of the computer device.
[0015] Figure 1 A flowchart of an information processing method provided in an embodiment of this application is shown below. Figure 1 As shown, the information processing method includes the following steps: Step S101: In response to the material selection operation, determine at least one target material content in the currently displayed interface.
[0016] The currently displayed interface can include the window interface of the information carrier or functional module that the user is currently operating on and viewing, such as a file, text component, table component, presentation component, email, webpage, etc. Material selection can be an operation where the user selects a range of material content of interest by clicking, highlighting, or selecting from screenshots. This material selection operation can be performed multiple times, with each selection corresponding to a specific segment of target material content. This target material content can be text content, image content, table content, etc. Optionally, the material selection operation can also be an operation where the user selects files of various formats by clicking or uploading through a specific interface. For example, documents, tables, presentations, audio, video, etc.
[0017] In one embodiment, for a document, its outline hierarchy information can be extracted first. This outline hierarchy information can be the hierarchical structure of the content in the document, divided according to logical relationships and importance. For example, different levels of headings or entries can describe the overall framework of the document and the subordinate relationships between its content. Taking multi-level headings as an example, a heading at a certain level can be selected to split the document content into multiple segments of target material content. Specifically, the range of the number of target material contents corresponding to the document can be determined based on the file size, and the level closest to that range can be selected from the multi-level headings.
[0018] In one embodiment, a table can be split according to data categories or a fixed number of rows. For example, a starting column can be selected and multiple category fields can be extracted. The table can then be split into multiple sub-tables according to these category fields, and each sub-table can be considered a segment of target content. Alternatively, the table can be split into multiple sub-tables according to a fixed number of rows, and each sub-table can be considered a segment of target content.
[0019] In one embodiment, for audio, the audio can be divided into multiple sub-audio segments according to a preset time granularity, and speech recognition can be performed on each sub-audio segment to obtain a corresponding sub-text. Then, the similarity of the sub-texts corresponding to adjacent sub-audio segments is calculated. If the similarity reaches a preset threshold, the sub-texts corresponding to the adjacent sub-audio segments can be merged to obtain a new sub-text. Finally, each sub-text obtained after similarity merging can be regarded as a segment of target material content.
[0020] In one embodiment, for a video, audio can be extracted from the video, and the audio can be processed according to the aforementioned embodiments to obtain at least one segment of target material content. Then, the video is processed by frame extraction according to a preset sampling ratio to obtain multiple image frames, each of which can be regarded as a segment of target material content. Optionally, the material selection operation can also be a user's copy operation for text, images, etc. Specifically, at least one copied content corresponding to the user's triggered copy operation can be obtained, the at least one copied content can be deduplicated, and then each remaining copied content can be determined as a segment of target material content. Of course, pairwise similarity calculation can also be performed on the at least one copied content. If the similarity reaches a preset threshold, the corresponding copied content can be merged to obtain new copied content. Then, each copied content obtained after similarity merging can be regarded as a segment of target material content.
[0021] In one embodiment, Figure 2 This application provides an embodiment of a schematic diagram illustrating the method of determining target material content through word selection and bounding, as shown below. Figure 2As shown, by receiving the user's word selection operation in the document area 201, the text content 202 can be determined, and the text content 202 is determined as the first target material content through the first reference control 203. Correspondingly, the first material control 205 corresponding to the first target material content is displayed in the dialog area 204. The dialog area 204 is used to receive user input information and feedback information processing results.
[0022] In one embodiment, Figure 3 This application provides an embodiment of a diagram illustrating how to determine the content of target material by selecting a frame in a screenshot. Figure 3 As shown, by receiving the user's screenshot selection operation in the table area 301, the image content 302 can be determined, and the image content 302 is determined as the second target material content through the second reference control 303. Correspondingly, the second material control 304 corresponding to the determined second target material content is displayed in the dialog area 204. The dialog area 204 can display multiple selected target material contents.
[0023] In one embodiment, Figure 4 A flowchart illustrating a process for determining target material content is provided in this application embodiment, such as... Figure 4 As shown, the specific steps to determine at least one segment of target material content in the current display interface are as follows: Step S1011: Extract information from the selected area in the current display interface to obtain the first material content.
[0024] The selected area can be the region defined and selected by the material selection operation, specifically a rectangle, an irregular polygon, a lasso outline, etc., which are not limited in this application. The first material content can be relevant information extracted from the area defined by the selected area, such as text content, image content, table content, etc.
[0025] In one embodiment, since the selected area is a partial content extraction of a portion of the currently displayed interface, there may be anomalies such as truncated sentences or table cells in the text or table content. Therefore, the obtained first source content can be inspected to determine if the aforementioned anomalies exist and appropriate completion can be made to ensure content coherence. Specifically, if the first source content is text content, punctuation and grammatical structure feature detection can be performed on it. Punctuation feature detection can check whether there are terminating punctuation marks (such as periods, question marks, exclamation marks, etc.) at the end of the sentence. If the content ends with a comma, semicolon, or colon, it can be judged as an abnormal truncation. Furthermore, it can also check whether the punctuation matches. If the content contains a left quotation mark but lacks a right quotation mark, it can be judged as an abnormal truncation. Grammatical structure feature detection can check whether sentence components are missing, such as subject and predicate. If either is missing, it can be considered an abnormal truncation. It can also check whether it only contains modifiers such as attributives or adverbs. If so, it can be considered as an abnormal truncation. After completing punctuation and grammatical structure feature detection, if an anomalously truncated sentence is detected, text extraction can be performed on adjacent areas based on the region containing that sentence until the sentence passes the punctuation or grammatical structure feature detection. Furthermore, if the first source material is table content, cell integrity detection can be performed to determine if each cell within the table has a closed outer border. If a cell's outer border only contains a portion of the border (e.g., only the top and left borders are displayed, lacking the bottom and right borders), or if the border lines are clearly truncated (e.g., the right border abruptly ends without a closed edge), then the cell can be determined to be truncated, potentially indicating missing internal information. After completing cell integrity detection, if an anomalously truncated cell is detected, border extension can be performed on adjacent areas based on the border of that cell until the internal information and border of the cell are complete.
[0026] Step S1012: Determine the associated information from the associated location area corresponding to the selected area.
[0027] The associated location region can be another region that is spatially or logically related to the selected region, but has not been directly selected by the user. In one embodiment, the specific implementation process for determining the associated information from the associated location region corresponding to the selected region is as follows: Locate the corresponding associated location area based on the information source object corresponding to the selected area; identify the associated location area based on the content type corresponding to the first material content to obtain the associated information.
[0028] The information source object can be a specific information carrier or functional module from which the selected area originates, such as a document, image, chart, email, or webpage. In one embodiment, if the information source object is a document, the associated location area can be the document title area, adjacent paragraph areas, etc., and the content type corresponding to the first material content is text. In this case, the title can be extracted from the document title area, and the context paragraphs can be extracted from the adjacent paragraph areas. The title and context paragraphs can be considered as associated information.
[0029] In one embodiment, if the information source object is an image, the associated location region can be several pixel regions adjacent to the selected region. Since the content type corresponding to the first material content is an image, these pixel regions can be identified and extracted to obtain several target entity sub-images. These target entity sub-images can be considered as associated information. Therefore, background information related to the selected region can be accurately extracted.
[0030] In one embodiment, if the information source object is a chart, the associated location area can be an axis area, a data label area, etc. For line charts, scatter plots, etc., the associated location area can be an axis area, used to define the data dimensions and numerical range of the chart, specifically including axis labels, data units, axis scale values, etc. For donut charts, pie charts, etc., the associated location area can be a data label area, used to label the specific information of each data point, specifically including classification, percentage, and numerical value, etc. Since the content type corresponding to the first material content is a graphic, the axis area and data label area can be identified and extracted to obtain the axis information and data label information, which can be considered as associated information.
[0031] In one embodiment, if the information source object is a table, the associated location area can be a row field area, a column field area, etc., and the content type corresponding to the first material content is table content, then the target row field information and target column field information can be obtained by identifying and extracting the row field area and column field area according to the row and column where the first material content is located. The target row field information and target column field information can be regarded as associated information.
[0032] Step S1013: Merge the first material content and related information corresponding to the same selected area to obtain the target material content.
[0033] Specifically, the first material content and related information can be merged according to their content types to obtain the target material content.
[0034] In one embodiment, if the content type of the first material content corresponding to a certain selected area is text, and the extracted associated information is a title and context paragraphs, then the context paragraphs can be added to the top and bottom of the first material content respectively, and marked with dashed boxes or other markers to distinguish them as associated information of the first material content. Then, the title is added to the front and marked with dashed boxes or other markers to obtain the target material content corresponding to the selected area.
[0035] In one embodiment, if the content type of the first material content corresponding to a certain selected area is an image, and the extracted association information is several entity sub-images, then the several entity sub-images can be spliced with the first material content according to the position of their original pixel areas relative to the selected area, and dotted boxes and other markers can be added to distinguish them as association information of the first material content, thereby obtaining the target material content corresponding to the selected area.
[0036] In one embodiment, if the content type of the first material content corresponding to a selected area is a chart, and the extracted associated information is coordinate axis information and data label information, then the coordinate axis information and data label information can be added to the adjacent area of the chart, and a dashed box or other markers can be added to obtain the target material content corresponding to the selected area.
[0037] In one embodiment, if the content type of the first material content corresponding to a selected area is table content, and the extracted associated information is target row field information and target column field information, the target row field information and target column field information can be added to the beginning of the relevant row or column according to the original row and column correspondence, and marked with dashed boxes, etc., to obtain the target material content corresponding to the selected area.
[0038] Therefore, by determining the associated information from the associated location area corresponding to the selected area, the background information related to the selected area can be effectively captured. By merging the first material content and associated information corresponding to the same selected area, the original material content can be supplemented with contextual information, thereby improving the semantic integrity of the material content.
[0039] Step S102: Obtain user input information, and determine at least one target fragment object from at least one target material content based on the user input information.
[0040] The user input information can describe the specific reasoning task that the AI model needs to process. Since at least one piece of target content may contain redundant information unrelated to the user input, it is necessary to filter out target fragments strongly related to the user input from this target content to improve the accuracy and efficiency of subsequent model reasoning. Target fragments can include content segments from the target content that are relevant to the user input.
[0041] In one embodiment, if the target content consists of a single segment, redundant information can be removed from this segment, and the corresponding metadata can be used to determine the target fragment object. In another embodiment, if the target content consists of multiple segments, target content matching the modality type mentioned in the user input information can be selected from these segments. This modality type can be text, images, tables, etc. Then, the target content is segmented, and the corresponding metadata is used to obtain multiple target fragment objects. In yet another embodiment, if the target content consists of multiple segments, these segments can be further segmented, and the corresponding metadata can be used to obtain multiple candidate fragment objects. These candidate fragment objects are then matched with the user input information based on similarity to determine the final target fragment objects.
[0042] Optionally, before determining at least one target fragment object from at least one segment of target material content based on user input information, the following implementation process is also included: The system performs keyword recognition on user input to determine the target modality type. If there is no target content corresponding to the target modality type in at least one segment of target content, the system extracts the keyword information corresponding to the target modality type from the user input, generates supplementary fragment information based on the keyword information, and displays the supplementary fragment information on the current display interface.
[0043] Different target modality types correspond to different keywords. For example, if the user input contains words like "paragraph," "translation," or "summary," it can be considered that the associated target modality type is text. Similarly, if the user input contains words like "image" or "screenshot," it can be considered that the associated target modality type is image. And if the user input contains words like "cell" or "table," it can be considered that the associated target modality type is table. Of course, there are other target modality types and corresponding matching keywords, which are not limited here. If the target material content in at least one segment does not contain target material content corresponding to the target modality type—for example, if the target modality type corresponding to the user input includes both text and image, but the target material content in at least one segment only contains target material content with the target modality type of text—then there will be a problem of missing material content, which will affect the accuracy of the subsequent inference results of the artificial intelligence model. Therefore, keyword information corresponding to the target modality type can be extracted from the user input information to generate supplementary information for the segment. For example, if the target material content of the target modality type (image) is missing in the aforementioned example, then it is necessary to obtain the keyword information corresponding to the image from the keyword recognition results corresponding to the user input information. For example, if the keyword information is "XXX screenshot", it means that the current target material content is missing a XXX screenshot. Therefore, the corresponding supplementary information "Please provide a XXX screenshot" can be generated and displayed in the dialog box area of the current display interface, or as a pop-up window, to prompt the user to promptly select the XXX screenshot. This application does not impose any limitations on this. Thus, by combining the multimodal information of the user input information, the system can promptly remind the user to supplement the missing target material content, ensuring the integrity of subsequent model input data and improving the accuracy of inference results.
[0044] In one embodiment, before generating and displaying supplementary fragment information based on keyword information in the current display interface, a search can be performed on the content referenced by the user within a preset time range based on the keyword information to determine if there is any related content. If related content is found, it is directly added as the target content without the need for subsequent generation and display of supplementary fragment information. Specifically, each content in the content library can be configured with descriptive information. By calculating the similarity between the keyword information and the descriptive information, the content with the highest similarity and reaching a preset threshold can be selected as the target content. In another embodiment, before generating and displaying supplementary fragment information based on keyword information in the current display interface, basic information such as file name and content title can be extracted from at least one open file in the current display interface. It can be determined whether the basic information of each file contains the aforementioned keyword information. If the basic information of a file contains the aforementioned keyword information, the file can be segmented to obtain the target content to be supplemented without the need for subsequent generation and display of supplementary fragment information.
[0045] Step S103: Determine prompt word information based on user input information and at least one target fragment object, and input the prompt word information into the set artificial intelligence model to obtain information processing results.
[0046] The prompt information can be an input instruction constructed by combining user input information and at least one target fragment object to optimize the understanding of the artificial intelligence model.
[0047] For example, the target content identified in the current display interface includes the following two sections: "Sales statistics table for various products in Region A + remarks: This month, the total sales in Region A were 4.5 million, a year-on-year increase of 12.3%, with high-end products contributing over 60%" from a spreadsheet file; and "...Last month, we added two new models to the high-end product line in Region A and adjusted our channel promotion strategy to increase agent enthusiasm. Customer feedback indicates high acceptance of new high-end products in Region A, with frequent stockouts in some stores. We recommend increasing inventory and optimizing delivery schedules...". The user input could be, "Please summarize the sales situation of high-end products in Region A." Based on the user input information, the following target segments can be identified from the aforementioned two target material contents: "Sales statistics table of high-end products in Region A [from the first sub-table of the table file]", "This month's total sales in Region A were 4.5 million, a year-on-year increase of 12.3%, of which high-end products contributed over 60% [from the remarks information in the table file]", "Last month we added two new models to the high-end product line in Region A and adjusted the channel promotion strategy to improve the enthusiasm of agents [from the weekly report document]", and "Customer feedback indicates high acceptance of new high-end products in Region A, with frequent stockouts in some stores; it is recommended to increase inventory and optimize delivery schedules [from the weekly report document]". A total of four target segments can be generated based on the user input information and the aforementioned target segments: "Based on the following information snippets, summarize the sales performance of high-end products in Region A. The output format must be: I. Core Background II. Key Content III. Conclusions and Recommendations." Requirements: 1. Identify the differences in theme, structure, tone, and details, and integrate them into a coherent new semantic whole. 2. Focus on causal reasoning and analyze the causal relationships between different snippets. 3. Integrate different perspectives to generate multi-dimensional explanations. 4. Convert the style and tone of different snippets into a unified "report style." The snippets are: "High-end product sales statistics table for Region A [from the first sub-table of the table file]", "This month, the total sales in Region A were 4.5 million, a year-on-year increase of 12.3%, of which high-end products contributed over 60% [from the notes in the table file]", "Last month, we added two new models to the high-end product line in Region A and adjusted the channel promotion strategy to improve the enthusiasm of agents [from the weekly report document]", and "Customer feedback indicates that Region A has a high acceptance of new high-end products, and some stores frequently experience stockouts. We recommend increasing inventory and optimizing the delivery schedule [from the weekly report document]". After inputting the prompt word information into the artificial intelligence model, the information processing results are as follows: "I. Core Background: The data in this summary comes from the sales statistics table of high-end products in Region A and related weekly reports and table notes. The statistical period focuses on this month. In order to improve the sales performance of high-end products and the enthusiasm of agents, two new models were added to the high-end product line in Region A last month, and the channel promotion strategy was adjusted at the same time, which laid the strategic foundation for this month's sales."
[0048] II. Key Takeaways: In terms of sales performance, total sales in Region A reached 4.5 million yuan this month, a year-on-year increase of 12.3%. High-end products contributed over 60% of this growth, becoming the core driver of overall sales growth. It is speculated that the addition of new product models and adjustments to channel strategies last month played a positive role in promoting performance growth. Regarding market feedback and operational status, customer feedback indicates a high acceptance of new high-end products in Region A. However, some stores experienced stockouts, reflecting a mismatch between current inventory levels and delivery schedules, indicating insufficient supply-demand coordination.
[0049] III. Conclusion and Recommendations: Overall, last month's product and channel strategies were effective, contributing to the sales growth of high-end products. However, shortcomings in inventory and delivery need to be addressed. Recommendations include: first, based on market demand forecasts, increasing inventory levels in key stores for high-end products, especially new ones; second, optimizing delivery schedules and establishing a dynamic inventory monitoring mechanism to ensure stable supply to stores, preventing stockouts from impacting sales conversion and customer experience, and further consolidating the role of high-end products in driving overall performance. In one embodiment, if there is only one target fragment object, the user input information and the target fragment object can be directly concatenated, or the prompt word information can be obtained by filling in a preset task template that matches the user input information. The preset task template can be a structured instruction pre-designed for a specific task type (e.g., translation, summary generation, comparative analysis, question-and-answer response, creation, etc.). This preset task template contains a fixed text portion and placeholders to be filled in. These placeholders are used to fill in the user input information and the target fragment object to generate prompt word information that matches the task scenario. Specifically, the user input information can first undergo task intent recognition, which can include explicit keyword recognition and implicit keyword recognition. Explicit keyword recognition can be performed through word matching using a preset thesaurus, which can store keywords for various tasks, such as "translation," "analysis," and "summary." Implicit keyword recognition can extract semantic associations from the user input information through a pre-trained semantic model. For example, if the user inputs "I want to know what this content is about," the semantic model can associate it with "summary." Then, after obtaining the explicit and implicit keywords, the template library can be queried based on these keywords to obtain the corresponding preset task templates. This template library can be categorized by task type, storing different preset task templates, such as creative tasks (including emails, copywriting, poetry, and stories), processing tasks (including translation, summarization, and format conversion), and analysis tasks (including data interpretation, problem diagnosis, and advantage comparison).
[0050] In one embodiment, if there are multiple target fragment objects, the vector similarity between each target fragment object and the user input information can be calculated first. The target fragment objects are then sorted from highest to lowest similarity value. The user input information and the sorted target fragment objects are then entered into a preset task template to obtain prompt word information. Optionally, before entering the sorted target fragment objects into the preset task template, the similarity value corresponding to each target fragment object can be converted into a reference weight according to a preset numerical range. The reference weight corresponding to each target fragment is added to the beginning or end of the target fragment, and then the processed target fragment objects are entered into the preset task template. Optionally, before entering the sorted target fragment objects into the preset task template, segmentation tags can be added to each target fragment object to help the model distinguish different target fragment objects. This artificial intelligence model can be a large language model, which is not limited here. Optionally, a historical material library can be set for the current user. This historical material library can collect historical material content referenced by the current user in the past. Each historical material content can be set with corresponding content description information or content description vector. In one embodiment, each historical content in the historical content library has a corresponding content description vector. User input information can be converted into an intent vector using models such as Sentence-BERT or ColBERT. This intent vector reflects the semantic features of the user's intent. The similarity between this intent vector and the content description vector corresponding to each historical content is calculated, yielding a similarity score for each content description. All content descriptions in the historical content library are sorted from highest to lowest similarity, and a predetermined number of content descriptions are selected. The historical content corresponding to these predetermined number of content descriptions is considered to have the highest relevance to the user input and can be recalled. In another embodiment, each historical content in the historical content library has a corresponding content description. At least one keyword can be extracted from the user input. It is determined whether each content description matches the aforementioned keyword or has similar keywords. The historical content with the most matches or the most similar keywords is considered to have the highest relevance to the user input and can be recalled. After recalling at least one relevant historical material, the recalled historical material can be added to the prompt word information, and the supplemented prompt word information can be input into the set artificial intelligence model to obtain the information processing result.
[0051] Optionally, after inputting the prompt word information into the set artificial intelligence model to obtain the information processing result, the following implementation process is also included: The dialog box area in the current display interface displays the information processing results and generates the fragment control corresponding to the target fragment object associated with the information processing results; when the fragment control is triggered, the partial content information and original source information of the target fragment object are displayed.
[0052] The dialog box area can be a pre-set or dynamically generated interactive area in the current display interface, used to receive user input, display the information processing results of the artificial intelligence model, and display other related content. The information processing results can be the structured output obtained by the artificial intelligence model after reasoning and analyzing based on prompt information. These results can include multiple text paragraphs, each referencing different target fragment objects. For example, if the number of target fragment objects input to the artificial intelligence model is N, and each target fragment object has a corresponding number, then if a text paragraph is generated using target fragment objects numbered 2 and 5, these target fragment objects can be considered associated with that text paragraph. Correspondingly, fragment controls corresponding to target fragment objects numbered 2 and 5 can be generated at the beginning, end, or middle of the text paragraph. These fragment controls can be visual interactive elements associated with the target fragment object, used to display partial content information and original source information of the target fragment object when triggered. This partial content information can be obtained by extracting partial content from the target fragment object. For example, if the target fragment object is a piece of text, the partial content information can be a portion of the text content, such as the first sentence or multiple keywords within the first sentence. For example, if the target fragment object is an image, the local content information can be a local image region of a preset size within that image, such as a local image region located in the upper left corner, a local image region located in the center, etc. The original source information can be related to the file, component, webpage, etc., from which the target fragment object originated. For example, if the target fragment object is a piece of text, and its corresponding original source is a document, then the original source information can be the document name; if its corresponding original source is a webpage, then the original source information can be the webpage link. Similarly, if the target fragment object is an image, and its corresponding original source is a table, then the original source information can be the table name; if its corresponding original source is a picture, then the original source information can be the picture name.
[0053] Figure 5 This application provides a schematic diagram illustrating how to trigger a fragment control in a dialog box area, as shown in the embodiment of the present application. Figure 5As shown, the dialog box area 204 of the current display interface displays the information processing result 501 and generates a fragment control 502. When the fragment control 502 is in the triggered state, partial content information is displayed in the first area 503 and the original source information is displayed in the second area 504.
[0054] As described above, in response to the material selection operation, identifying at least one target material content in the current display interface allows users to quickly select the required material content through intuitive interface operations, accurately locating the material range. Based on user input information, identifying at least one target fragment object from at least one target material content can accurately filter material fragments matching the user's intent. Determining prompt word information based on user input information and at least one target fragment object, and inputting the prompt word information into the set artificial intelligence model to obtain information processing results, can effectively integrate fragmented information and improve the accuracy of the model's output. This solution can flexibly call upon different material contents selected by the user for cross-modal information integration, effectively match the user's actual intent for information reasoning, meet the user's specific personalized task needs, and improve the accuracy and reliability of the reasoning results.
[0055] Figure 6 A flowchart of an information processing method including a process for determining a target fragment object is provided for an embodiment of this application, as shown below. Figure 6 As shown, the information processing method includes the following steps: Step S601: In response to the material selection operation, determine at least one target material content in the currently displayed interface.
[0056] Step S602: Obtain user input information and parse at least one piece of target material content to determine multiple candidate fragment objects.
[0057] Since the target content originates from unstructured multimodal data directly selected by the user, it may suffer from issues such as information redundancy and noise interference. Therefore, the target content can be parsed and segmented into structured data, which can be candidate fragment objects with relatively independent semantics.
[0058] In one embodiment, Figure 7 A flowchart illustrating a process for determining candidate fragment objects is provided in this application embodiment, as follows: Figure 7 As shown, the specific steps for parsing at least one piece of target material to determine multiple candidate fragment objects are as follows: Step S6021: Identify the material format of each target material content segment, and perform semantic parsing on each target material content segment corresponding to its material format to obtain structured semantic information.
[0059] The source material can be in formats such as text, images, and tables. By identifying the format of each target source material, accurate semantic analysis can be performed using specific grammatical rules or recognition methods corresponding to that format. For example, if the target source material is in text format, semantic analysis needs to be performed based on preset grammatical rules and hierarchical relationships. As another example, if the target source material is in image format, it can be identified first using optical character recognition or edge detection. If text or tables are present, they are extracted. Then, semantic analysis can be performed on text based on preset grammatical rules and hierarchical relationships, on tables based on row and column relationships and structural features, and on images using object detection models (such as YOLO, Faster R-CNN, etc.). This structured semantic information can be structured data obtained by standardizing and transforming the target source material according to its corresponding hierarchical structure, type features, and element relationships.
[0060] Step S6022: Clean and segment the structured semantic information corresponding to each segment of target material content to obtain multiple candidate content fragments.
[0061] Cleaning structured semantic information can be a process of modifying, filtering, and standardizing it, aiming to remove noise, correct errors, and unify formatting. For example, removing HTML (Hypertext Markup Language) or rich text tags, invalid characters, and style noise. Another example is unifying the layout or format of text, tables, and images. Segmenting target content can involve breaking down structured semantic information into smaller, relatively independent semantic units to obtain candidate content fragments. For example, text content can be broken down into titles and paragraphs; similarly, screenshot content can be broken down into text, images, and tables.
[0062] Step S6023: Integrate each candidate content fragment and its corresponding metadata to obtain multiple candidate fragment objects.
[0063] Meta-information can record the source attributes of candidate content fragments. Specifically, this meta-information can include basic information, location information, etc. Basic information can include filename, file path, creation time, author, etc., while location information can include page number, worksheet name, component name, etc. In one embodiment, the meta-information corresponding to each candidate content fragment can be combined with the candidate content fragment in key-value pair form or with additional fields to obtain a candidate fragment object. Optionally, each candidate fragment object can also have fields such as structure identifier and modality type. Taking the content type of the candidate fragment object as text as an example, the structure identifier can be the hierarchical structure to which the candidate fragment object belongs, such as which level of heading, which paragraph, etc., while the modality type is text. Taking the content type of the candidate fragment object as a table as an example, the structure identifier can be the row and column area to which the candidate fragment object belongs, while the modality type is table. Thus, at least one piece of target material content can be standardized and output to obtain multiple candidate fragment objects, which is beneficial for subsequent rapid retrieval and location of target fragment objects, improving model inference efficiency. In one embodiment, the specific implementation process of integrating each candidate content fragment and its corresponding meta-information to obtain multiple candidate fragment objects is as follows: Each candidate content fragment and its corresponding meta-information are combined to obtain a target content fragment; each target content fragment is converted into a corresponding semantic embedding vector and at least one cluster center is calculated; based on the at least one cluster center, each target content fragment is grouped and merged to obtain multiple candidate fragment objects.
[0064] This process involves converting each target content fragment into a corresponding semantic embedding vector. Specifically, this can be achieved using a set semantic embedding model, such as the Sentence-BERT model or the ColBERT model, which is not limited in this application. The cluster center can be a vector representing the core features of the cluster, serving as the center of all semantic embedding vectors within the cluster and used to measure the correlation between other fragments and the cluster. In one embodiment, a predetermined number of cluster centers can be randomly selected from the semantic embedding vectors based on a preset number of clusters. The distance from each semantic embedding vector to the cluster center is then calculated, and the cluster to which the semantic embedding vector belongs is reassigned. Then, the cluster centers are recalculated based on the semantic embedding vectors within the cluster until the cluster centers no longer change significantly or the predetermined number of iterations is reached. Finally, the target content fragments corresponding to all semantic feature vectors within each cluster can be merged to obtain candidate fragment objects. In one embodiment, the distance and density between all semantic embedding vectors can be calculated first. Based on the obtained distance and density, it can be identified whether different semantic embedding vectors belong to the same density connected region, and all semantic embedding vectors within the same density connected region can be merged to form a cluster. Finally, the target content fragments corresponding to all semantic feature vectors within each cluster can be merged to obtain candidate fragment objects.
[0065] Step S603: Based on user input information, perform similarity retrieval on multiple candidate fragment objects to determine at least one target fragment object.
[0066] After identifying multiple candidate fragment objects, user input information can be matched against these candidate fragment objects to determine the target fragment object with high similarity. This similarity retrieval can be performed by calculating vector cosine similarity or by keyword matching; this application does not specify a particular method.
[0067] In one embodiment, Figure 8 A flowchart illustrating a process for retrieving and determining a target fragment object is provided in this application embodiment, as follows: Figure 8 As shown, the specific steps for determining at least one target fragment object by performing similarity retrieval on multiple candidate fragment objects based on user input information are as follows: Step S6031: Vectorize the user input information to obtain a first vector, and vectorize multiple candidate fragment objects to obtain a second vector set.
[0068] The first vector can be obtained by using lightweight LLM components or intent classification models, such as T5, DistilBERT, and FastText, to vectorize user input information. The second vector set can be obtained by using models such as Sentence-BERT, ColBERT, and CLIP to vectorize each candidate fragment object.
[0069] Step S6032: Perform similarity calculations between the first vector and each second vector in the second vector set to obtain the similarity score.
[0070] The similarity calculation can be performed using cosine similarity or Euclidean distance, etc., and this application does not limit the calculation. The higher the calculated similarity, the higher the similarity between the first vector and the second vector.
[0071] Step S6033: Determine the candidate fragment objects corresponding to the second vectors in the second vector set that meet the similarity conditions as the target fragment objects.
[0072] Since the number of candidate fragment objects may be large, to reduce the amount of data processed by the model, second vectors that meet the similarity criteria can be selected from the second vector set. In one embodiment, determining the second vectors that meet the similarity criteria can be done by sorting multiple second vectors in the second vector set from high to low similarity and selecting a preset number of second vectors at the top of the sorted list. For example, if the preset number is 5, then the top 5 second vectors in the similarity ranking will be selected. In another embodiment, determining the second vectors that meet the similarity criteria can be done by selecting second vectors from the second vector set whose similarity is greater than a preset threshold. The selected second vectors correspond to the candidate fragment objects most relevant to the user input information; therefore, these candidate fragment objects can be identified as the target fragment objects.
[0073] This allows for the accurate filtering of contextual information that is more relevant to user input, enabling cross-modal retrieval and providing precise input data for artificial intelligence models.
[0074] Step S604: Determine prompt word information based on user input information and at least one target fragment object, and input the prompt word information into the set artificial intelligence model to obtain information processing results.
[0075] As described above, by parsing at least one piece of target material content to determine multiple candidate fragment objects, the target material content can be transformed into standardized output. Based on user input information, similarity retrieval is performed on multiple candidate fragment objects to determine at least one target fragment object. This allows for rapid matching of associated fragment objects, providing multimodal content with stronger semantic associations for subsequent model inference.
[0076] Figure 9 A flowchart of an information processing method including a process for determining prompt word information is provided for an embodiment of this application, such as... Figure 9 As shown, the information processing method includes the following steps: Step S901: In response to the material selection operation, determine at least one target material content in the currently displayed interface.
[0077] Step S902: Obtain user input information, and determine at least one target fragment object from at least one target material content based on the user input information.
[0078] Step S903: Determine the target task template from the multiple candidate task templates set according to the user input information.
[0079] The candidate task template can be a predefined prompt word framework to design corresponding structured instructions for different task types. For example, candidate task templates could be "summary generation," "comparative analysis," or "question and answer response." By performing semantic similarity matching or keyword matching on user input information, the target task template can be determined from multiple candidate task templates.
[0080] Step S904: Sort and integrate multiple target fragment objects to obtain a fragment object sequence.
[0081] The sorting and integration process may include one or more processes such as merging overlapping segments, sorting according to preset rules, and content selection. The sorting according to preset rules may be based on the generation timestamp or the similarity, which is not limited in this application.
[0082] In one embodiment, the specific implementation process of sorting and integrating multiple target fragment objects to obtain a fragment object sequence is as follows: Sort multiple target fragment objects according to their similarity to user input information from high to low; then truncate and merge the sorted target fragment objects according to a preset window length to obtain a fragment object sequence.
[0083] In cases where there are multiple target fragment objects, to ensure that the AI model prioritizes target fragment objects with high relevance to the user input information, these multiple target fragment objects can be sorted from high to low similarity to the user input information, thus strengthening the priority processing of key information. The preset window length can be the product of the preset window percentage corresponding to each target fragment object and the context window length defined by the AI model, used to control the total amount of fragment content finally input into the AI model, ensuring the stability of model inference. Specifically, for the sorted multiple target fragment objects, the first few complete target fragment objects can be truncated and retained according to the preset window length, and then merged to obtain a fragment object sequence.
[0084] Step S905: Fill the user input information and fragment object sequence into the target task template to obtain prompt word information.
[0085] Step S906: Input the prompt word information into the set artificial intelligence model to obtain the information processing result.
[0086] As mentioned above, matching the target task template corresponding to the user input information can help the artificial intelligence model understand the task requirements and reduce the threshold for user input. Sorting and integrating multiple target fragment objects to obtain a fragment object sequence can organize multiple target fragment objects in an orderly manner, avoid simply piling up target fragment objects, and improve the model's reasoning effect.
[0087] Figure 10 This is a structural block diagram of an information processing apparatus provided in an embodiment of this application. The apparatus is configured to execute the information processing method provided in the above embodiments, and possesses corresponding functional modules and beneficial effects for executing the method. For example... Figure 10 As shown, the device specifically includes: The material content determination module 1001 is configured to determine at least one piece of target material content in the currently displayed interface in response to the material selection operation; The target segment determination module 1002 is configured to acquire user input information and determine at least one target segment object from at least one piece of target material content based on the user input information; The processing result determination module 1003 is configured to determine prompt word information based on user input information and at least one target fragment object, and input the prompt word information into the set artificial intelligence model to obtain the information processing result.
[0088] As described above, in response to the material selection operation, identifying at least one target material content in the current display interface allows users to quickly select the required material content through intuitive interface operations, accurately locating the material range. Based on user input information, identifying at least one target fragment object from at least one target material content can accurately filter material fragments matching the user's intent. Determining prompt word information based on user input information and at least one target fragment object, and inputting the prompt word information into the set artificial intelligence model to obtain information processing results, can effectively integrate fragmented information and improve the accuracy of the model's output. This solution can flexibly call upon different material contents selected by the user for cross-modal information integration, effectively match the user's actual intent for information reasoning, meet the user's specific personalized task needs, and improve the accuracy and reliability of the reasoning results.
[0089] In one possible embodiment, the target fragment determination module 1002 is further configured to: Analyze at least one segment of target material to identify multiple candidate segment objects; Based on user input, a similarity search is performed on multiple candidate fragment objects to determine at least one target fragment object.
[0090] In one possible embodiment, the target fragment determination module 1002 is further configured to: Identify the format of each target material content segment, and perform semantic parsing on each target material content segment corresponding to its format to obtain structured semantic information; The structured semantic information corresponding to each segment of target material is cleaned and segmented to obtain multiple candidate content fragments; The candidate content fragments and their corresponding metadata are integrated to obtain multiple candidate fragment objects.
[0091] In one possible embodiment, the target fragment determination module 1002 is further configured to: The user input information is vectorized to obtain the first vector, and multiple candidate fragment objects are vectorized to obtain the second vector set; The similarity is calculated by comparing the first vector with each of the second vectors in the second vector set; The candidate fragment objects corresponding to the second vectors in the second vector set that meet the similarity conditions are determined as the target fragment objects.
[0092] In one possible embodiment, the processing result determination module 1003 is further configured to: The target task template is determined from multiple candidate task templates based on user input. Multiple target fragment objects are sorted and integrated to obtain a sequence of fragment objects; The user input information and the sequence of fragment objects are filled into the target task template to obtain prompt information.
[0093] In one possible embodiment, the processing result determination module 1003 is further configured to: Sort multiple target fragment objects in descending order of their similarity to user input information; The sorted target fragment objects are truncated and merged according to a preset window length to obtain a fragment object sequence.
[0094] In one possible embodiment, the material content determination module 1001 is further configured as follows: Extract information from the selected area in the current display interface to obtain the first material content; Determine the associated information from the associated location area corresponding to the selected area; The target content is obtained by merging the first material content and related information corresponding to the same selected area.
[0095] In one possible embodiment, the material content determination module 1001 is further configured as follows: Locate the corresponding associated location area based on the information source object corresponding to the selected area; The associated information is obtained by identifying the associated location area based on the content type corresponding to the first material.
[0096] In one possible embodiment, a fragment supplementation suggestion module is also included, configured as follows: Keyword recognition is performed on user input to determine the target modality type; If at least one piece of target content does not contain target content corresponding to the target modality type, extract the keyword information corresponding to the target modality type from the user input information, generate supplementary fragment information based on the keyword information, and display the supplementary fragment information on the current display interface.
[0097] In one possible embodiment, a fragment reference reverse lookup module is also included, configured as follows: Display the information processing results in the dialog area of the current display interface, and generate the fragment control corresponding to the target fragment object associated with the information processing results; When a trigger fragment control is detected, display partial content information and original source information of the target fragment object.
[0098] Figure 11 This is a schematic diagram of the structure of an information processing device provided in an embodiment of this application, such as... Figure 11As shown, the device includes a processor 1101, a memory 1102, an input device 1103, and an output device 1104; the number of processors 1101 in the device can be one or more. Figure 11 Taking a processor 1101 as an example; the processor 1101, memory 1102, input device 1103, and output device 1104 in the device can be connected via a bus or other means. Figure 11 Taking a bus connection as an example, the memory 1102, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the information processing method in this embodiment. The processor 1101 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 1102, thereby realizing the aforementioned information processing method. The input device 1103 can be configured to receive input digital or character information and generate key signal inputs related to user settings and function control of the device. The output device 1104 may include a display device such as a screen.
[0099] The information processing device provided above can be used to execute the information processing method provided in any of the above embodiments, and has corresponding functions and beneficial effects.
[0100] This application also provides a non-volatile storage medium containing computer-executable instructions, which, when executed by a computer processor, are configured to perform an information processing method described in the above embodiments, comprising: in response to a material selection operation, determining at least one segment of target material content in the current display interface; acquiring user input information, determining at least one target fragment object from the at least one segment of target material content based on the user input information; determining prompt word information based on the user input information and the at least one target fragment object, and inputting the prompt word information into a set artificial intelligence model to obtain an information processing result.
[0101] Storage medium – any type of memory device or storage device. The term “storage medium” is intended to include: mounting media, such as CD-ROMs, floppy disks, or magnetic tape devices; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media, optical storage; registers or other similar types of memory elements, etc. Storage media may also include other types of memory or combinations thereof. Furthermore, storage media may reside in a first computer system in which the program is executed, or may reside in a different second computer system connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term “storage medium” can include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). Storage media may store program instructions (e.g., specifically implemented as a computer program) executable by one or more processors.
[0102] Of course, the computer-executable instructions provided in the embodiments of this application are not limited to the information processing method described above, but can also perform related operations in the information processing method provided in any embodiment of this application.
[0103] It should be noted that the numbering of each step in this solution is only used to describe the overall design framework of this solution and does not indicate a necessary sequential relationship between the steps. As long as the overall implementation process conforms to the overall design framework of this solution, it falls within the protection scope of this solution. The literal order in the description is not an exclusive limitation on the specific implementation process of this solution. Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent memory in computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0104] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0105] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. An information processing method characterized by comprising: The method comprises the following steps: In response to a material selection operation, at least one piece of target material content is determined in a current display interface; User input information is obtained, and at least one target segment object is determined from the at least one piece of target material content based on the user input information; Prompt word information is determined according to the user input information and the at least one target segment object, and the prompt word information is input into a set artificial intelligence model to obtain an information processing result.
2. The information processing method according to claim 1, characterized by, The at least one target segment object is determined from the at least one piece of target material content based on the user input information, comprising: The at least one piece of target material content is parsed to determine a plurality of candidate segment objects; Similarity retrieval is performed on the plurality of candidate segment objects based on the user input information to determine at least one target segment object.
3. The information processing method according to claim 2, characterized by, The at least one piece of target material content is parsed to determine a plurality of candidate segment objects, comprising: The format of each piece of target material content is identified, and semantic parsing is performed on each piece of target material content according to the format of the material to obtain structured semantic information; The structured semantic information corresponding to each piece of target material content is cleaned and segmented to obtain a plurality of candidate content segments; Each candidate content segment and corresponding meta information are integrated to obtain a plurality of candidate segment objects.
4. The information processing method according to claim 3, characterized by, The at least one target segment object is determined from the at least one piece of target material content based on the user input information, comprising: Each candidate content segment and corresponding meta information are combined to obtain a target content segment; Each target content segment is converted into a corresponding semantic embedding vector, and at least one cluster center is calculated; Based on the at least one cluster center, each target content segment is grouped and combined to obtain a plurality of candidate segment objects.
5. The information processing method according to claim 2, characterized by, The at least one target segment object is determined from the at least one piece of target material content based on the user input information, comprising: The user input information is vectorized to obtain a first vector, and the plurality of candidate segment objects are vectorized to obtain a second vector set; Similarity calculation is performed between the first vector and each second vector in the second vector set to obtain a similarity; The second vector in the second vector set that meets the similarity condition is determined as a target segment object.
6. The information processing method according to claim 1, characterized by, The number of target segment objects is a plurality, and the prompt word information is determined according to the user input information and the at least one target segment object, comprising: A target task template is determined from a plurality of candidate task templates set according to the user input information; A segment object sequence is obtained by sorting and integrating a plurality of target segment objects; The user input information and the segment object sequence are filled into the target task template to obtain prompt word information.
7. The information processing method according to claim 6, characterized by, The plurality of target segment objects are sorted according to the similarity with the user input information from high to low; The sorted plurality of target segment objects are content-truncated and merged according to a preset window length to obtain a segment object sequence. 8. The information processing method according to claim 1, characterized by, The at least one piece of target material content in the current display interface includes: The information extraction of the selected region in the current display interface obtains the first material content; The associated information in the associated position region corresponding to the selected region is determined; The first material content and the associated information corresponding to the same selected region are merged to obtain the target material content.
9. The information processing method according to claim 8, characterized by, The associated information in the associated position region corresponding to the selected region includes: The corresponding associated position region is positioned according to the information source object corresponding to the selected region; The associated position region is identified based on the content type corresponding to the first material content to obtain the associated information.
10. The information processing method according to claim 1, characterized by, Before the at least one target segment object is determined from the at least one piece of target material content based on the user input information, the method further includes: The target modality type is determined by keyword recognition of the user input information; In the case that the target material content corresponding to the target modality type does not exist in the at least one piece of target material content, the keyword information corresponding to the target modality type is extracted from the user input information, the segment supplement information is generated based on the keyword information, and the segment supplement information is displayed in the current display interface.
11. The information processing method according to claim 1, characterized by, After the information processing result is obtained by inputting the prompt word information into the set artificial intelligence model, the method further includes: The information processing result is displayed in the dialog box region in the current display interface, and the segment control corresponding to the target segment object associated with the information processing result is generated; In the case that the segment control is triggered, the local content information and the original source information of the target segment object are displayed.
12. An information processing apparatus comprising: The method includes: The material content determination module is configured to determine at least one piece of target material content in the current display interface in response to a material selection operation; The target segment determination module is configured to obtain user input information, and determine at least one target segment object from the at least one piece of target material content based on the user input information; The processing result determination module is configured to determine prompt word information according to the user input information and the at least one target segment object, and input the prompt word information into a set artificial intelligence model to obtain an information processing result.