A query processing method, an electronic device and a storage medium based on a large model, relating to artificial intelligence fields such as deep learning, large models, intelligent cloud, natural language processing and retrieval-augmented generation, are provided. The method may include: acquiring a target query to be processed; retrieving based on the target query from M predetermined data sources to obtain a first retrieval result, and inputting the target query and the first retrieval result into a target large model, where M is a positive integer; acquiring and displaying a target answer output by the target large model in a streaming output manner, where the target answer is an answer corresponding to the target query, and during a streaming output process, performing keyword recognition on output text content, and generating and displaying a knowledge card corresponding to a keyword.
Provided is a method for identifying a hotspot external link, an electronic device, and a storage medium, relating to the field of data processing technologies, and especially to the fields offle sharing and data statistics. The method includes: in response to an operation of a user on an external link, obtaining an external link identifier, an area code of the user and an operation type, to construct a triple counter, wherein each triple counter corresponds to a dimension of (external link, area, operation type); performing hierarchical processing and filtering on the triple counter via multi-level cache, to identify a potential hotspot triple; identifying a sustained hotspot external link based on the potential hotspot triple by incorporating counter values of a current time period and previous M historical time periods, wherein M is a positive integer; and outputting an identification result of the hotspot external link.
Provided is a method for generating subtitle, an electronic device and a storage medium, relating to the field of computer technologies, and especially to the technical fields of artificial intelligence, natural language processing, signal processing, software engineering and interactive design. The method includes: determining a target video; segmenting audio data of the target video to obtain a plurality of target audio segments; obtaining a plurality of text segments by performing speech recognition on the plurality of target audio segments according to a concurrent processing mechanism, wherein the plurality of text segments correspond to the plurality of target audio segments in a one-to-one correspondence; obtaining a target subtitle text by performing translation processing on the plurality of text segments.
G06F 40/58 - Use of machine translation, e.g. for multi-lingual retrieval, for server-side translation for client devices or for real-time translation
G10L 17/02 - Preprocessing operations, e.g. segment selectionPattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal componentsFeature selection or extraction
G10L 25/57 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination for processing of video signals
4.
LIVE STREAM SWITCHING METHOD, ELECTRONIC DEVICE AND STORAGE MEDIUM
A live stream switching method and an electronic device are provided. The method includes: receiving a stream switching request sent by a client, where the stream switching request includes a target resolution; pulling an original live stream from a first source and a target live stream corresponding to the target resolution from a second source, and transmitting the original live stream to the client; performing timestamp alignment between the original live stream and the target live stream; and in response to determining that alignment is successful, ceasing transmission of the original live stream and transmitting the target live stream to the client.
H04N 21/238 - Interfacing the downstream path of the transmission network, e.g. adapting the transmission rate of a video stream to network bandwidthProcessing of multiplex streams
Provided is a method for processing a user query, an electronic device and a storage medium, relating to the field of computer technology, and in particular to the fields of artificial intelligence, large language model and other technologies. The method includes: receiving a first user query; performing word segmentation on the first user query to obtain a token list corresponding to the first user query; determining a reusable key-value cache of the first user query according to the token list; and when the reusable key-value cache is in a process of transmission from a GPU to a CPU, stopping the transmission of the reusable key-value cache, and allocating the reusable key-value cache to the first user query.
Provided is a video generation method, electronic device and storage medium, which relates to the field of artificial intelligence technology, and in particular to the fields of computer vision, deep learning, large models, and augmented reality. The method includes: extracting a reference feature of a target object from a reference image; determining a mixed pose sequence; wherein a plurality of target parts in the mixed pose sequence adopt a corresponding description mode to enhance a pose characteristic of a corresponding part; introducing initial noise to a denoising network, and performing a reverse denoising operation on the reference feature and the mixed pose sequence, to generate a video of the target object.
G06V 10/77 - Processing image or video features in feature spacesArrangements for image or video recognition or understanding using pattern recognition or machine learning using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]Blind source separation
G06V 40/10 - Human or animal bodies, e.g. vehicle occupants or pedestriansBody parts, e.g. hands
A scheduling method for a model inference request is provided. The implementation includes: determining, based on a model inference request to be scheduled, at least one first data item to be encoded; determining a cache hit rate corresponding to each model instance based on a first data index and the at least one first data item, where the first data index includes: a plurality of second data items determined by performing deduplication processing based on a plurality of historical data items cached by the plurality of model instances; and a data identifier corresponding to each second data item, which includes a plurality of sub-identifiers, each sub-identifier indicating whether the model instance corresponding to the sub-identifier caches the second data item; determining a target model instance based on the cache hit rate corresponding to each model instance; and scheduling the model inference request to the target model instance to perform inference.
Beijing Baidu Netcom Science Technology Co Ltd. (China)
Inventor
Zhang, Yuqin
Ma, Yanjun
Yu, Dianhai
Shen, Liang
Abstract
Provided is a hybrid pipeline scheduling method for distributed deep learning training, an electronic device and a storage medium, relating to the technical field of artificial intelligence, and particularly to the technical field of distributed training, machine learning and deep training. The method includes: dividing training data into N consecutive chunks, wherein each chunk includes K micro-batches; performing calculations on first N−1 chunks by using an interleaved forward-then-backward scheduling strategy; and performing calculations on an N-th chunk by using an interleaved 1F1B scheduling strategy. Among them, the performing calculations on the N-th chunk by using the interleaved 1F1B scheduling strategy includes: for any micro-batch in the N-th chunk, triggering a backward calculation immediately upon completion of its forward calculation; and for any micro-batch, releasing a video memory occupied by an activation value generated during its forward calculation immediately upon completion of its backward calculation.
A method for generating codes based on artificial intelligence (AI) includes: obtaining block text data of a multimedia text and a problem text to be processed, in which the block text data includes a plurality of block texts and block information of each block text, and the block information includes formula information; selecting target block texts from the plurality of block texts based on the plurality of block texts, the formula information in each block information and the problem text; and performing code generation processing based on the target block texts, block information of the target block texts and the problem text.
Provided is a method for question answering over list data based on a large model, an electronic device and a storage medium, relating to the field of data processing technology, and especially to the field of natural language processing. The method includes: obtaining a question text; processing a known data list based on the question text using a large model, to generate structured associated data; performing optimization processing on the structured associated data, to generate a knowledge base collection; inputting the question text and the knowledge base collection into the large model, to enable the large model to generate an answer corresponding to the question text. The technical solution can enhance the model's focus on data, significantly improve the accuracy and efficiency of question answering over lists.
Provided is a method for encoding a video, an electronic device and a storage medium, relating to the field of video encoding. The method includes: obtaining a reference frame that is adjacent to a current frame and has established hierarchical division of an initial ROI; reusing a motion vector generated by the reference frame in a motion compensated temporal filter process to determine an ROI position prediction value of the current frame; extracting facial feature points from the current frame, and correcting the ROI position prediction value based on a spatial topological relationship of feature points in the initial ROI; configuring differentiated quantization parameters for a main face region, a secondary face region and a background region after correction respectively; dynamically allocating a three-region bitrate according to a real-time network bandwidth; and outputting an encoded frame of the current frame and associated ROI-level metadata.
H04N 19/167 - Position within a video image, e.g. region of interest [ROI]
G06V 10/44 - Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersectionsConnectivity analysis, e.g. of connected components
G06V 20/40 - ScenesScene-specific elements in video content
G06V 40/16 - Human faces, e.g. facial parts, sketches or expressions
H04N 19/105 - Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
H04N 19/172 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a picture, frame or field
H04N 19/196 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the adaptation method, adaptation tool or adaptation type used for the adaptive coding being specially adapted for the computation of encoding parameters, e.g. by averaging previously computed encoding parameters
A method for video understanding, an apparatus, an electronic device, and a storage medium are disclosed, which relates to artificial intelligence fields such as deep learning, large models, computer vision and natural language processing. The method includes: sampling a video to be processed to obtain M sampled original images, wherein M is a positive integer greater than 1; obtaining a text recognition result of audio corresponding to the video to be processed; determining a target input information based on the respective original images and the text recognition result; inputting the target input information into a video understanding model to obtain a video understanding result corresponding to the video to be processed.
G06V 20/40 - ScenesScene-specific elements in video content
G06V 10/62 - Extraction of image or video features relating to a temporal dimension, e.g. time-based feature extractionPattern tracking
G06V 10/75 - Organisation of the matching processes, e.g. simultaneous or sequential comparisons of image or video featuresCoarse-fine approaches, e.g. multi-scale approachesImage or video pattern matchingProximity measures in feature spaces using context analysisSelection of dictionaries
G06V 20/62 - Text, e.g. of license plates, overlay texts or captions on TV images
G10L 15/18 - Speech classification or search using natural language modelling
G10L 25/57 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination for processing of video signals
13.
METHOD FOR HYBRID THINKING MODEL DISTILLATION, ELECTRONIC DEVICE, AND STORAGE MEDIUM
A method for hybrid thinking model distillation, including: adjusting, based on training progress, a data ratio between thinking mode data and non-thinking mode data in training data; obtaining, based on adjusted training data, a first output sequence generated by a teacher model in the non-thinking mode and a second output sequence generated by the teacher model in the thinking mode; obtaining, based on the adjusted training data, a third output sequence generated by a student model in the non-thinking mode and a fourth output sequence generated by the student model in the thinking mode; and training the student model based on the first output sequence, the second output sequence, the third output sequence, and the fourth output sequence.
The present disclosure provides a disaster recovery processing method and a disaster recovery processing apparatus for CDN live streaming, a device, and a medium, which relates to the field of cloud computing technologies, and in particular, to CDN (content delivery network) live streaming field. A specific implementation is: a streaming quality analysis platform acquires quality data of an edge node and a central node from a system log, and determines a fault node and a fault type thereof using a fault identification model or threshold ranges of the quality data. The identified fault node information is sent to a node configuration platform, to generate configuration blocks including filters and processing. The configuration blocks are issued to respective nodes. The node determines, based on streaming information, whether a filter is hit, and performs a corresponding disaster recovery processing when it is hit.
H04L 41/0654 - Management of faults, events, alarms or notifications using network fault recovery
H04L 41/16 - Arrangements for maintenance, administration or management of data switching networks, e.g. of packet switching networks using machine learning or artificial intelligence
H04L 65/61 - Network streaming of media packets for supporting one-way streaming services, e.g. Internet radio
15.
MULTI-MODAL-FUSED METHOD AND APPARATUS FOR RECOGNIZING HIGH-DEFINITION MAP ELEMENT, AND DEVICE AND MEDIUM
Provided are a multimodal fusion-based high-definition map feature recognition method and apparatus, a device, and a medium. The method includes determining attribute characteristics, pixel registration characteristics, and hybrid registration characteristics of a target map feature based on point cloud data and at least two types of candidate image data of the target map feature; determining a pixel correspondence of the target map feature based on the pixel registration characteristics of the target map feature and determining a hybrid correspondence of the target map feature based on the hybrid registration characteristics of the target map feature; and fusing the attribute characteristics of the target map feature based on the pixel correspondence and the hybrid correspondence to obtain a fused characteristic of the target map feature and determining a feature category of the target map feature based on the fused characteristic of the target map feature.
G06V 10/764 - Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
An image processing method, an apparatus, a device, and a medium are provided, which relate to the technical field of image processing, specifically to the technical field of image denoising, image enhancement and the like. The image processing method comprises: obtaining an image to be processed, where the image to be processed comprises a plurality of pixels; determining a surface curvature of each pixel of the plurality of pixels; determining, based on a first predetermined curvature requirement, an isolated point noise pixel in the plurality of pixels; and denoising the isolated point noise pixel.
H04N 19/86 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using pre-processing or post-processing specially adapted for video compression involving reduction of coding artifacts, e.g. of blockiness
17.
CAMERA IMAGE DATA PROCESSING METHOD, APPARATUS, ELECTRONIC DEVICE, AND MEDIUM
A method includes: obtaining initial camera parameters; determining, based on the initial camera parameters, a target condition; determining, based on the target condition, at least one first time offset, where each of the at least one first time offset indicates an exposure delay time between every two adjacent cameras of a plurality of cameras, and when setting the exposure delay time between every two adjacent cameras of the plurality of cameras based on each first time offset, the first traffic peak generated during the transmission of a plurality of images in parallel to a processor by the plurality of cameras is less than a first threshold; and setting, based on the at least one first time offset, the exposure delay time between every two adjacent cameras of the plurality of cameras.
H04N 23/698 - Control of cameras or camera modules for achieving an enlarged field of view, e.g. panoramic image capture
H04N 23/80 - Camera processing pipelinesComponents thereof
H04N 23/90 - Arrangement of cameras or camera modules, e.g. multiple cameras in TV studios or sports stadiums
H04N 25/40 - Extracting pixel data from image sensors by controlling scanning circuits, e.g. by modifying the number of pixels sampled or to be sampled
Provided are a method for automatic parallelization of a mixture of experts model, an apparatus for automatic parallelization of a mixture of experts model, a device, a medium, and a program. The method includes acquiring a computational graph of the mixture of experts model; determining a process mesh of an expert weight tensor of an expert model, where the processes are supported to execute by computing devices in the distributed system; splitting a global data tensor into sub-data tensors required by corresponding expert models, and configuring a process mesh of a sub-data tensor to be the same as a process mesh of a corresponding expert model; performing, by the expert model, processing based on an input sub-data tensor and the expert weight tensor to output a sub-result tensor; and determining a result tensor of the mixture of experts model based on at least one sub-result tensor.
A model training snapshot backup method, performed by a first training node, is provided. The method includes: obtaining dynamic parameters to be backed up of a current training step of a model; determining a training snapshot of the current training step according to the dynamic parameters to be backed up; and backing up the training snapshot to a memory of the first training node.
G06F 11/14 - Error detection or correction of the data by redundancy in operation, e.g. by using different operation sequences leading to the same result
20.
CODE KNOWLEDGE GRAPH GENERATION METHOD AND APPARATUS, CODE GENERATION METHOD AND APPARATUS, DEVICE, AND MEDIUM
A code knowledge graph generation method and apparatus, a code generation method and apparatus, a device, and a medium, relate to the field of data processing, specifically the technical fields of intelligent search, human-computer interaction, artificial intelligence, and large language models. The specific implementation solution includes acquiring a code tree, where the code tree includes first code elements and a structural relationship among the first code elements, where the code tree is generated by performing content parsing on a source code; generating at least one first graph node according to the first code elements; generating a first graph edge between corresponding graph nodes according to the structural relationship; generating a code knowledge graph according to the at least one first graph node and the first graph edge.
Method and apparatus for training multimodal large model and method and apparatus for image question answering are disclosed, which relates to artificial intelligence technologies such as large models, deep learning, natural language processing, and computer vision. The method for training multimodal large model includes: obtaining an initial sample image, a sample object in the initial sample image, and a location information of the sample object; obtaining a target sample image including a sample visual marker based on the initial sample image and a target image region corresponding to the initial sample image; obtaining a sample question corresponding to the target sample image based on the sample visual marker, and obtaining a sample answer corresponding to the sample question; training an initial multimodal large model based on a target training sample constituted by the target sample image, the sample question and the sample answer to obtain a target multimodal large model. The method for image question answering includes: obtaining a target image including a target visual marker and a target question; inputting the target image and the target question into the target multimodal large model to obtain a target answer. The present disclosure enables the target multimodal large model to effectively understand the target visual marker in the target image, thereby improving the accuracy of the target answer.
G06V 10/22 - Image preprocessing by selection of a specific region containing or referencing a patternLocating or processing of specific regions to guide the detection or recognition
Method and apparatus for text processing based on large model, and method and apparatus for large model compression are disclosed, which relate to the technical field of artificial intelligence field such as deep learning, large model, and natural language processing. The method for text processing based on large model includes: obtaining a token sequence corresponding to an input text; performing the following processing respectively for respective tokens in the token sequence: in response to determining that a fusion layer in a target large model needs to be used to process a token, generating a target processing result corresponding to the token by executing inference computation in the fusion layer at least twice, wherein the target large model is obtained by performing a model compression on a large model to be compressed, the model compression includes fusing Lm consecutive layers in the large model to be compressed into the fusion layer, Lm is a positive integer greater than 1, and Lm is less than L, where L represents the number of layers included in the large model to be compressed.
A training method for a large model and a data processing method is provided, relating to the technical field of data processing, and particularly to artificial intelligence and large model technologies. The implementation is: determining a training sample set, where the training sample includes a pair of sample input and sample output, and the sample output includes a natural language output and a code block, where the natural language output includes a code execution result generated by the code block; providing the sample input to a large model to obtain a predicted output of the large model; and adjusting, based on the predicted output and the sample output, parameters of the large model.
A training method for a neural network model for image processing is provided. The present disclosure relates to the technical field of artificial intelligence, and in particular to the technical field of image recognition. The neural network model includes a first sub-model and a second sub-model, and a training method for the first sub-model includes: obtaining a first sample image and labeling a ground truth coordinate value of a region of interest in the first sample image; inputting the first sample image into the first sub-model to obtain a first output; and adjusting parameters of the first sub-model; a training method for the second sub-model includes: obtaining a second sample image and labeling a ground truth threshold; inputting the second sample image into the first sub-model and obtaining a second output of the first sub-model; inputting the second output into the second sub-model and obtaining a predicted threshold output by the second sub-model; and adjusting parameters of the second sub-model based on the ground truth threshold and the predicted threshold.
G06V 10/28 - Quantising the image, e.g. histogram thresholding for discrimination between background and foreground patterns
G06V 10/25 - Determination of region of interest [ROI] or a volume of interest [VOI]
G06V 10/26 - Segmentation of patterns in the image fieldCutting or merging of image elements to establish the pattern region, e.g. clustering-based techniquesDetection of occlusion
G06V 10/77 - Processing image or video features in feature spacesArrangements for image or video recognition or understanding using pattern recognition or machine learning using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]Blind source separation
G06V 10/774 - Generating sets of training patternsBootstrap methods, e.g. bagging or boosting
G06V 10/82 - Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
A method of recognizing a speech, a device, and a medium. The method includes: processing, by using an acoustic model, speech data to be recognized and a first text segment obtained by recognition to obtain respective acoustic probabilities of a plurality of candidate text segments; processing the first text segment by using a first language sub-model to obtain respective initial language probabilities of the plurality of candidate text segments; processing the first text segment by using a constraint sub-model to obtain extendibility relationships of the plurality of candidate text segments with respect to the first text segment; adjusting the initial language probabilities of the candidate text segments according to the extendibility relationships to obtain respective first language probabilities of the plurality of candidate text segments; and determining a target text segment from the plurality of candidate text segments according to the first language probabilities and the acoustic probabilities.
A method is provided for scheduling a video transcoding task, and relates to the field of data processing technology, and in particular to artificial intelligence and streaming media technology. The implementation is: determining an input video to be transcoded, wherein the input video includes at least one image frame; performing an image analysis on the at least one image frame to obtain an image feature of the input video; determining target a transcoding parameter for a transcoded target video; predicting, based on the image feature of the input video and the target transcoding parameter, a resource occupancy required for a transcoding task for the input video and an acceptable memory usage for the transcoding task; scheduling a computational resource for the transcoding task based on the predicted resource occupancy and memory usage.
H04N 19/40 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using video transcoding, i.e. partial or full decoding of a coded input stream followed by re-encoding of the decoded output stream
H04N 19/127 - Prioritisation of hardware or computational resources
H04N 19/156 - Availability of hardware or computational resources, e.g. encoding based on power-saving criteria
H04N 19/423 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by implementation details or hardware specially adapted for video compression or decompression, e.g. dedicated software implementation characterised by memory arrangements
H04N 19/436 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by implementation details or hardware specially adapted for video compression or decompression, e.g. dedicated software implementation using parallelised computational arrangements
27.
INFERENCE ACCELERATION METHOD AND ELECTRONIC DEVICE FOR LARGE MODELS
An inference acceleration method relating to artificial intelligence technical fields such as a large model, deep learning, and natural language processing is provided. The inference acceleration method for large models includes: after inputting a source text to be processed into a target large model, obtaining a top-layer hidden state of the target large model for predicting a next token; obtaining action decision information corresponding to the next token according to the top-layer hidden state; in response to determining that the action decision information is a copy action, obtaining a text copy interval corresponding to the next token according to the top-layer hidden state; copying text in the source text to be processed that is located within the text copy interval, and using a copy result as the next token.
A human-computer interaction method, a human-computer interaction apparatus, an electronic device and a storage medium are provided, which relate to a field of artificial intelligence technology, and in particular to fields of deep learning, natural language processing and large model technologies. The human-computer interaction method includes: determining, in response to a human-computer interaction request, a first target plug-in related to a first dialogue text contained in the human-computer interaction request from a plurality of plug-ins registered in a large language model based on the first dialogue text contained in the human-computer interaction request; obtaining a second dialogue text based on the first dialogue text and a description text of the first target plug-in; and inputting the second dialogue text into the large language model to obtain a response text.
A method and apparatus for generating a long text, an electronic device, a computer readable storage medium are provided. An embodiment of the method includes: generating a long text outline based on long text requirement information, the long text outline including a chapter entry; generating, in response to receiving file data associated with the chapter entry, a text fragment corresponding to the chapter entry based on the file data; and generating the long text based on the long text outline and the text fragment corresponding to the chapter entry.
A method for interacting voice is provided. The method includes: determining a user included in a physical environment and a first position of the user in the physical environment based on a real-time audio stream collected in the physical environment; presenting a user indicator corresponding to the user in association with a target indicator in a voice interaction interface rendered for the physical environment, where the relative positional relationship between the user indicator and the target indicator is determined based on the relative positional relationship between the first position and a second position corresponding to the target indicator in the physical environment; and adjusting a visual presentation attribute of the user indicator based on a portion of the real-time audio stream corresponding to the user.
A data generation method, a model training method, a data processing method, an electronic device, and a storage medium are provided, which relate to the field of artificial intelligence technologies, and in particular to the fields of large model and intelligent agent technologies. The data generation method includes: generating at least one response result to be evaluated using at least one of data synthesis expert units; determining at least one evaluation result for the at least one response result to be evaluated using at least one of the data synthesis expert units; in response to determining that the at least one evaluation result indicates a presence of at least one response result to be corrected, determining at least one correction data according to at least one evaluation result for the at least one response result to be corrected; and returning to the generating step.
A method for generating a multimodal text, a method for acquiring a multimodal text, a device, and a medium are provided, which relate to the field of artificial intelligence technology, and in particular to technical fields of computer vision, deep learning and large models. The method for generating a multimodal text includes the follows: a text information corresponding to a prompt information is generated by a large language model based on the prompt information, in response to a multimodal text generation request including the prompt information being received; an image information corresponding to the text information is generated by the large language model based on the text information; and a multimodal text rendering tool is called by the large language model based on the text information and the image information to render the multimodal text including the text information and the image information.
A method of generating speech data based on a large model and a method of training a large model are provided, which relate to artificial intelligence technology, in particular to fields of speech generation, intelligent customer service, video production, etc. The method includes: receiving prosodic description text and speech text, where the prosodic description text describes pronunciation prosodic intentions for text characters in the speech text; performing semantic fusion on the prosodic description text and the speech text using the large model, to obtain a prosodic fusion feature, where a sub-feature in the prosodic fusion feature characterizes a pronunciation prosody of a speech segment with respect to the text characters; and generating, based on the prosodic fusion feature and a specified pronunciation attribute associated with a specified object, target speech data characterizing that the specified object pronounces in accordance with the pronunciation prosodic intentions and corresponding to the speech text.
A method for adaptive code processing based on artificial intelligence, including: obtaining a candidate code set, wherein the candidate code set comprises a plurality of candidate codes, first description information corresponding to each candidate code, and second description information corresponding to each code slice in the candidate codes; determining a similarity between every two candidate codes in the candidate code set based on a plurality pieces of first description information and a plurality pieces of second description information; obtaining a new candidate code set by performing crossover and mutation on two candidate codes with a similarity less than a first threshold; and returning to the step of determining the similarity based on the new candidate code set until a target code set meeting a requirement is obtained.
A method for training a large language model includes: determining a second operation procedure of a first sample text through a first large language model; obtaining a second webpage address obtained through an interaction between the first large language model and a sample browser based on the second operation procedure; determining a target reward value obtained through the interaction between the first large language model and the sample browser according to the second webpage address and a first webpage address corresponding to the first sample text; and performing a reinforcement learning training on the first large language model according to the target reward value.
An agent-based method for generating corpus data, an agent, a device, and a medium are provided which relate to the field of artificial intelligence technology, and in particular to the fields of deep learning, large models, and intelligent question answering technologies. The agent-based method for generating corpus data includes: acquiring a target task rule related to a page to be processed, where a semantic dependency relationship in the target task rule represents a semantic understanding logic among a plurality of item components in the page to be processed; executing a target task for the plurality of item components based on the semantic dependency relationship by using a designated agent according to the target task rule, and outputting an item response information; and generating target corpus data according to the item response information and item contents in the item components.
A method of processing a speech stream, which is related to the field of artificial intelligence technology, and more particularly to the fields of deep learning, speech processing, and voice conversion technologies, and includes: performing a feature extraction on a first speech frame sequence in a speech stream to be processed to obtain a first speech feature, where the first speech frame sequence overlaps with at least one second speech frame in a second speech frame sequence, and the second speech frame sequence precedes the first speech frame sequence in the speech stream; fusing, based on an attention mechanism, the first speech feature and a second speech feature determined based on the second speech frame sequence to obtain a speech fusion feature; and converting the speech fusion feature based on a preset speech attribute to obtain converted speech data corresponding to the first speech frame sequence.
G06N 3/0895 - Weakly supervised learning, e.g. semi-supervised or self-supervised learning
G10L 25/30 - Speech or voice analysis techniques not restricted to a single one of groups characterised by the analysis technique using neural networks
38.
METHOD FOR GENERATING ADAPTIVE PROGRAMS BASED ON ARTIFICIAL INTELLIGENCE, AGENT, AND STORAGE MEDIUM
A method for generating an adaptive program based on artificial intelligence (AI) performed by an agent in an electronic device. The method includes: determining diversity information of a target population, wherein the diversity information indicates program diversity of the target population; determining a temperature parameter according to the diversity information, wherein the temperature parameter is used to adjust a selection pressure; selecting, from the target population according to the temperature parameter, a first parent program and a second parent program corresponding to the first parent program; and obtaining a target program by performing an evolution iteration according to the first parent program and the second parent program.
Provided is a method for training a large model, an electronic device and a storage medium, relating to the field of computer technology, and in particular to the fields of data processing, deep learning, multimodal large model and other technologies. The method includes: inputting a query into a first large model so that the first large model responds to the query to obtain a response to the query; inputting the query and the response into the first large model so that the first large model evaluates and corrects the response to obtain an evaluation result and a correction result of the response; and fine-tuning the first large model based on at least one of the query, the response, the evaluation result or the correction result. The present disclosure can improve the accuracy and reliability of the response of the large model.
A method for updating a parameter of a large language model is provided. The method may include: generating target text using the large language model based on a text generation request; determining facts to be relied on in the generation process of the target text according to the text generation request to obtain a target fact set; determining reward data according to the target fact set and an information set including unverified information in the target text; and updating the parameter of the large language model based on the reward data.
Provided is a method for scheduling concurrent inference tasks, an electronic device and a storage medium, relating to the fields of artificial intelligence, deep learning, large model and other technologies. The method includes: determining multiple types of computing resources and a plurality of network models required for the concurrent inference tasks, wherein the concurrent inference tasks represent a plurality of inference tasks to be processed in parallel, and each network model is used to execute at least one of the plurality of inference tasks; determining actual execution time required for each model unit in the plurality of network models to execute a task on a candidate computing resource as well as actual resource switching time corresponding to each model unit, to obtain total execution time required to execute the concurrent inference tasks; and using the total execution time to determine a target scheduling result.
An image detection method includes: calling a teacher model to generate a first reconstructed image according to an initial image; calling a student model to generate a second reconstructed image according to the initial image, in which the student model is trained based on the teacher model and has a capability to detect an abnormal region in the image; and determining a position of the abnormal region in the initial image according to a reconstruction error between the initial image and the first reconstructed image and a reconstruction error between the initial image and the second reconstructed image.
A method of retrieving data, a method of training a deep learning model, an electronic device, and a storage medium are provided, which relate to a field of artificial intelligence technology, and in particular to fields of natural language processing and deep learning technologies. The method of retrieving data includes: determining M candidate texts from a text library based on a semantic information in a query to be processed, where M is an integer greater than or equal to 1; determining N candidate texts from the text library based on a keyword information in the query to be processed, where N is an integer greater than or equal to 1; and determining at least one target text based on the M candidate texts and the N candidate texts.
Provided is a driving information prediction method, apparatus, and autonomous driving vehicle, which relate to the field of autonomous driving, especially to the field of artificial intelligence, and particularly to the technical fields of autonomous driving and intelligent transportation. The method includes: determining a first leader-follower relationship between a target vehicle and a first obstacle based on a motion parameter of the target vehicle at a current moment, path information of the target vehicle within a first time period, a motion parameter of the first obstacle at the current moment, and predicted path information of the first obstacle within the first time period; obtaining first predicted driving information based on the first leader-follower relationship; and determining first optimal driving information based on an evaluation result corresponding to the first predicted driving information.
A method for predicting a structure of a compound includes: obtaining a combination of biomolecular sequences by combining specified biomolecular sequences; predicting a first probability distribution for the combination of biomolecular sequences, in which the first probability distribution is used for indicating first probabilities of candidate structural unit groups in the combination of biomolecular sequences, a candidate structural unit group includes structural units of at least two biomolecular sequences, and a first probability is used for indicating a possibility of interaction between structural units in a respective candidate structural unit group; determining, from the plurality of candidate structural unit groups, at least one first structural unit group based on the first probability distribution; and predicting a target structure of a biomolecular compound based on structural units interacted with each other in the at least one first structural unit group.
A task execution method, a large model training method, a device, and a medium are provided, which relate to the field of artificial intelligence technologies, and in particular to the fields of deep learning, computer vision, and large model technologies. The task execution method includes: acquiring an input information for executing a target task, where the input information includes a task description information and a current state information of a task object; inputting the task description information into a guidance large model to generate a guidance information; and inputting the current state information, the task description information, and the guidance information into a task execution agent to output a task execution result, where the guidance information is configured to guide the task execution agent to execute the target task on the task object.
A method for generating training data is performed by an electronic device, including: obtaining interaction data with an application, determining a state transition image based on the interaction data, generating a plurality of pieces of trajectory data based on the state transition image, obtaining reasoning reference information corresponding to the operation data in the trajectory data output by a multimodal model, by inputting the trajectory data into the multimodal model, and generating training data for training an interaction agent model based on the trajectory data and the reasoning reference information corresponding to the operation data in the trajectory data.
G06F 3/04845 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range for image manipulation, e.g. dragging, rotation, expansion or change of colour
50.
METHOD FOR TRAINING TEXT RETRIEVAL MODEL, METHOD FOR RETRIEVING TEXT, AND RELATED APPARATUSES
A method for training a text retrieval model, a method for retrieving a text and corresponding apparatuses are provided. An implementation of method comprises: compressing each pair of a query term sample and a candidate text sample into one character string sample; processing the character string sample using a feature extraction sub-model in a text retrieval model to obtain hidden-layer features of tokens in the character string sample; calculating feature weights of the hidden-layer features of the tokens using a weight determination sub-model in the text retrieval model; calculating, based on the feature weights and the hidden-layer features of the corresponding tokens, to obtain a correlation score between the query term sample and the candidate text sample using a similarity calculation sub-model in the text retrieval model; and training, based on the correlation score, the text retrieval model by means of contrastive learning.
An image coding method, which relates to the field of image processing technologies, and in particular to the fields of video compression and image encoding/decoding technologies, includes: partitioning an image to be processed into a plurality of coding units; determining, for any coding unit, a gradient histogram of the coding unit and respective gradient histograms of a plurality of sub-blocks of the coding unit, according to gradients of pixels in the coding unit; determining consistency between a dominant gradient direction of each of the plurality of sub-blocks and a dominant gradient direction of the coding unit according to the gradient histograms; determining a set of candidate coding modes from a plurality of angular coding modes according to the consistency; and determining a target angular coding mode to encode the coding unit, according to rate-distortion costs of the angular coding modes in the set of candidate coding modes.
H04N 19/11 - Selection of coding mode or of prediction mode among a plurality of spatial predictive coding modes
H04N 19/149 - Data rate or code amount at the encoder output by estimating the code amount by means of a model, e.g. mathematical model or statistical model
H04N 19/176 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
A lane-level positioning method and apparatus, a device, a vehicle, and a medium are provided. The method includes: obtaining map road information and visual perception road information of surroundings of a vehicle based on positioning information of the vehicle; determining a plurality of candidate lanes based on the positioning information; determining a topological recursion probability of the vehicle being in each of the candidate lanes by utilizing the map road information; determining a perception observation probability of the vehicle being in each of the candidate lanes by utilizing the visual perception road information; determining a positioning probability of the vehicle being in each of the candidate lanes by utilizing the positioning information; determining a target lane the vehicle is in from the plurality of candidate lanes based on the topological recursion probability, the perception observation probability, and the positioning probability of the vehicle being in each of the candidate lanes.
A method includes: obtaining a document comprising at least one page for question answering; determining a first vector corresponding to each of the at least one page; determining a second vector corresponding to a target question text to be answered; performing the following first operations: determining, based on the second vector and the first vector corresponding to each of the at least one page, a first similarity between the target question text and each of the at least one page; determining, based on the first similarities, at least one candidate page with the highest similarity to the target question text among the at least one page; and generating, based on the at least one candidate page and the target question text, a first identifier and first content, or second identifier and second content, using a large language model.
Provided is a method for converting animal language, an electronic device and a storage medium, relating to the field of artificial intelligence technology, and specifically to the fields of machine learning, deep learning, natural language processing and other technologies. The method includes: obtaining multimodal data related to an animal, wherein the multimodal data comprises animal sound data, animal behavior data and animal physical sign data; preprocessing the multimodal data to obtain fused multimodal data; recognizing current emotion of the animal according to the fused multimodal data to obtain an emotion recognition result of the animal; and performing semantic mapping and language translation on the emotion recognition result to convert animal language into human language to obtain a language conversion result.
G10L 25/63 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination for estimating an emotional state
55.
METHOD FOR AUTOMATICALLY ANNOTATING AN OBSTACLE, ELECTRONIC DEVICE AND STORAGE MEDIUM
Provided is a method for automatically annotating an obstacle, an electronic device and a storage medium, relating to the field of artificial intelligence technology, and in particular, to technologies fields of autonomous driving, neural network, deep learning and the like. The method includes: optimizing a target parameter in a projection relationship based on a re-projection error, the projection relationship is used to project a target obstacle from a reference frame onto a frame to be optimized, and satisfies a constraint in which positions of the target obstacle in different frames are consistent in an obstacle coordinate system established according to the target obstacle; and determining a target pose of the target obstacle based on the optimized target parameter.
G06V 20/58 - Recognition of moving objects or obstacles, e.g. vehicles or pedestriansRecognition of traffic objects, e.g. traffic signs, traffic lights or roads
G06T 7/174 - SegmentationEdge detection involving the use of two or more images
G06T 7/246 - Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
G06T 7/55 - Depth or shape recovery from multiple images
G06T 7/73 - Determining position or orientation of objects or cameras using feature-based methods
G06V 10/26 - Segmentation of patterns in the image fieldCutting or merging of image elements to establish the pattern region, e.g. clustering-based techniquesDetection of occlusion
G06V 10/98 - Detection or correction of errors, e.g. by rescanning the pattern or by human interventionEvaluation of the quality of the acquired patterns
G06V 20/70 - Labelling scene content, e.g. deriving syntactic or semantic representations
56.
Structured query language evaluation method, electronic device and storage medium
Provided is a structured query language evaluation method, an electronic device and a storage medium, relating to a technical field of data processing, and specifically to technical fields of large model and natural language processing. The method includes: obtaining a query question; generating a predicted structured query language based on the query question by using a large language model; in presence of a target structured query language, evaluating accuracy of the predicted structured query language based on the predicted structured query language and the target structured query language to obtain an evaluation result corresponding to the predicted structured query language; and in absence of the target structured query language, evaluating the accuracy of the predicted structured query language based on a semantic analysis result of the predicted structured query language to obtain the evaluation result corresponding to the predicted structured query language.
A retrieval-augmented query-and-answer method includes acquiring a to-be-answered query; performing character-level matching and screening between the to-be-answered query and offline-stored material knowledge points in a database to obtain preliminarily screened material knowledge points; performing semantic-level matching and screening between the to-be-answered query and the preliminarily screened material knowledge points to obtain a finely screened material knowledge point; acquiring a bound material slice from the database based on the finely screened material knowledge point; generating prompt information based on the bound material slice and the to-be-answered query; and inputting the prompt information into a query-and-answer large model, processing the prompt information, and generating an answer. This solution can accelerate the query-and-answer processing speed based on a large model and improve the accuracy of crosslingual document retrieval.
An information presentation method based on a large model, a device, and a medium, which relate to the field of data processing technologies, and in particular to the field of artificial intelligence technologies such as large models, natural language processing, and deep learning. The method includes: matching, in response to receiving a query question, the query question against a set of target example pairs corresponding to a query type of the query question to obtain at least one reference example pair, where the large model is configured to generate a query statement for the query question using the reference example pair; invoking the large model according to a prompt information to generate a target query statement, where the prompt information is obtained based on the query question and the at least one reference example pair; and presenting a query result obtained by executing the target query statement.
The present disclosure provides a method for generating document question-answering data, a training method, a generating apparatus, an electronic device, a computer-readable storage medium, and a computer program product. The method includes: extracting page content from page images in a document to obtain descriptive information corresponding to seed pages in the document; generating a reasoning chain corresponding to the seed pages by using a preset question-answering data generation model based on the descriptive information, question definitions of preset question types, and question-answering examples of the preset question types; and in response to the reasoning chain constituting a question-type reasoning chain corresponding to the preset question types, generating question-answering data corresponding to the preset question types by using the question-answering data generation model based on the question-type reasoning chain.
A method for compressing prompt information includes: obtaining first prompt information of a language model (LM); obtaining target length constraint information; and obtaining second prompt information of the LM by compressing the first prompt information based on the target length constraint information; in which a number of first tokens of the first prompt information is greater than a number of second tokens of the second prompt information, and semantics of the second prompt information is relevant to semantics of the first prompt information.
Provided is a method for generating a digital human, an intelligent agent, an electronic device and a storage medium, relating to the field of artificial intelligence technology, and particularly to the fields of computer vision, deep learning, large model, augmented reality and other technologies. The method includes: segmenting a target object from an image to be processed to obtain a target sub-image; selecting a digital human to be optimized that is compatible with the target object from a digital human set based on the target sub-image; generating clothing texture of the digital human to be optimized based on an appearance feature of the target object in the target sub-image; applying the clothing texture to the digital human to be optimized to obtain a target digital human; and driving the target digital human.
G06V 10/40 - Extraction of image or video features
G06V 10/764 - Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
62.
METHOD FOR EVALUATING POSE IN STANDING LONG JUMP, ELECTRONIC DEVICE, AND STORAGE MEDIUM
Provided is a method for evaluating pose in standing long jump, an electronic device, and a storage medium, relating to the field of computer vision applicable to scenarios of physical education, fitness testing, and training for adolescents. The method includes: generating, for a video frame in a standing long jump video of a target object, a body spatial position, a joint angle, a relative position parameter of a body part and a body moving velocity of the target object in the video frame; extracting key motion frames from the standing long jump video; obtaining a key motion evaluation result of the target object in a key motion frame; and determining a standing long jump pose evaluation result according to the key motion evaluation result of the target object.
A method includes: obtaining a target image including a face of a target object; performing facial keypoint extraction on the target image to obtain a first facial keypoint image; obtaining a first set of expression coefficients based on the first facial keypoint image and a preset set of expression bases; adjusting a corresponding expression coefficient in the first set of expression coefficients to obtain a second set of expression coefficients; obtaining a second facial keypoint image based on the second set of expression coefficients and the set of expression bases; and obtaining, based on the second facial keypoint image and the target image, a first image corresponding to the target image that has undergone a facial expression transformation.
Provided is a method apparatus for evaluating a project based on a large model, an electronic device, and a storage medium, relating to the field of computer technology, and in particular to fields of software and hardware project development, software and hardware project evaluation, machine learning, large model and other applications. The method includes: obtaining an evaluation intention for a target project; where the evaluation intention is used to request an evaluation of the target project under a specific evaluation indicator; obtaining an initial evaluation result for the target project under the specific evaluation indicator based on the evaluation intention using the large model; and obtaining a target evaluation result for the target project based on the initial evaluation result.
G06Q 10/063 - Operations research, analysis or management
G06Q 10/04 - Forecasting or optimisation specially adapted for administrative or management purposes, e.g. linear programming or "cutting stock problem"
A data processing method, an electronic device, and a medium are provided. The method includes: in response to receiving a data reading request, determining a first storage file and a second storage file according to an identifier of data to be read in the request, where a storage mode of the first storage file is different from that of the second storage file, and a data length of the second storage file is greater than that of the first storage file; determining an index result according to a position information of the data to be read and first index information of the first storage file; reading first sub-data from the first storage file when the index result indicates that the first storage file includes the first sub-data; and reading second sub-data from the second storage file according to the position information and index information of the second storage file.
Provided is a target model training method, a multimodal data processing method, and devices therefor, relating to the field of artificial intelligence technology, and in particular to the fields of computer vision, deep learning, large model and other technologies. The target model training method includes: inputting sample data into a preset model to obtain initial multimodal features of the sample data; and using the initial multimodal features and a preset noise feature to perform model training on N diffusion networks in the preset model to obtain a target model when parameters of an image-text encoder of the preset model are fixed.
G06V 10/26 - Segmentation of patterns in the image fieldCutting or merging of image elements to establish the pattern region, e.g. clustering-based techniquesDetection of occlusion
G06V 10/77 - Processing image or video features in feature spacesArrangements for image or video recognition or understanding using pattern recognition or machine learning using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]Blind source separation
67.
METHOD FOR CONTROLLING VIDEO MEMORY FOR MODEL TRAINING, ELECTRONIC DEVICE AND STORAGE MEDIUM
A method for controlling a video memory for model training, an electronic device and a storage medium are provided, relating to the field of artificial intelligence technology, and in particular to the fields of neural network, large model, training optimization and other technologies. The method includes: reconstructing a video memory space for one or more backward calculations during model training according to grouping information of parameter gradient information required for the one or more backward calculations; performing the one or more backward calculations to obtain one or more backward calculation results; storing the one or more backward calculation results into the video memory space reconstructed for the one or more backward calculations; and releasing the video memory space reconstructed for the one or more backward calculations.
The present disclosure provides a method and apparatus for processing code, a method and apparatus for training a code model, and a device, which relates to the field of artificial intelligence, and in particular, to the technical fields of software development, data processing, large models and computer communication. The method for processing code includes: acquiring target pasted code in response to a code paste event being monitored; performing dependency analysis for the target pasted code to obtain context-dependent data corresponding to the target pasted code; performing, according to the context-dependent data, code optimization for the target pasted code to obtain optimized code corresponding to the target pasted code; and outputting the optimized code. The present disclosure can greatly improve development efficiency, seamlessly integrate optimized code into the current code development environment, and improve code quality.
A method for video transcoding, an apparatus, an electronic device and a readable storage medium are suggested, which relates to the technical field of artificial intelligence including video processing, big data, and cloud services. The method for video transcoding includes: decoding an initial compressed video file to obtain an original compressed video file and a decoding dataset acquired during a decoding process; obtaining a target decoding sub-data from a decoding data corresponding to a frame to be encoded according to a position of a first coding unit in the frame to be encoded; performing a predictive coding on the first coding unit according to the target decoding sub-data to obtain a first prediction block of the frame to be encoded; obtaining a target compressed video file according to each frame to be encoded and the corresponding first prediction block.
H04N 19/40 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using video transcoding, i.e. partial or full decoding of a coded input stream followed by re-encoding of the decoded output stream
H04N 19/105 - Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
H04N 19/159 - Prediction type, e.g. intra-frame, inter-frame or bidirectional frame prediction
H04N 19/176 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
H04N 19/196 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the adaptation method, adaptation tool or adaptation type used for the adaptive coding being specially adapted for the computation of encoding parameters, e.g. by averaging previously computed encoding parameters
A method of audio generation based on a large language model is disclosed, which involves the fields of artificial intelligence such as large language models, natural language processing, deep learning, and audio generation. The method of audio generation based on a large language model comprises: acquiring a text to be processed; parsing the text to be processed using the large language model to obtain role information and emotional information corresponding to the text to be processed; obtaining a target reference text and a target reference audio according to the role information and the emotional information; and generating a target audio corresponding to the text to be processed according to the text to be processed, the target reference text, and the target reference audio.
G10L 13/027 - Concept to speech synthesisersGeneration of natural phrases from machine-based concepts
G10L 13/08 - Text analysis or generation of parameters for speech synthesis out of text, e.g. grapheme to phoneme translation, prosody generation or stress or intonation determination
71.
MULTI-AGENT-BASED INFORMATION PROCESSING METHOD, ELECTRONIC DEVICE AND STORAGE MEDIUM
A multi-agent-based information processing method includes: receiving an information processing request, in which the information processing request includes input information; inputting the input information to a first agent and obtaining output information of the first agent, in which the first agent determines one or more second agents from a set of agents based on the input information; and obtaining response information corresponding to the input information based on the output information of the first agent and the one or more second agents.
A method for generating information is provided. The method includes determining a task type of a target task; determining a task evaluation dimension corresponding to the target task according to the task prompt word and the task type of the target task; generating an evaluation result corresponding to the task evaluation dimension according to the task evaluation dimension and the task result, where the task result is generated by a large language model according to a target task and a task prompt word; and determining target information of the target task according to the task evaluation dimension and the evaluation result.
An image generation method, an apparatus, an electronic device and a storage medium are provided. The method includes: discretizing a target text to obtain a plurality of text tokens; obtaining a resolution sequence based on an initial resolution and a target resolution, wherein the resolution sequence comprises a plurality of resolutions, and a difference between two adjacent resolutions of the plurality of resolutions is a preset increment; generating image tokens corresponding respectively to the plurality of resolutions based on the plurality of text tokens and the resolution sequence; and fusing all the image tokens to obtain a target image corresponding to the target text.
The disclosure provides a method and an apparatus for processing an image, an electronic device, and a storage medium, which relates to the field of artificial intelligence technologies, and particularly to a technical field such as computer vision, deep learning, and large-scale models. The solution includes: obtaining an input content adapted to an image processing task, in which the input content includes at least one of: a first text token sequence, a first image token sequence, or an image-text fusion sequence; obtaining a joint feature representation including multimodal semantic information by performing cross-modal semantic modeling on the input content, in which the multimodal semantic information indicates a semantic correlation relationship of the input content in different modalities; and generating an output content adapted to the image processing task based on the joint feature representation.
G06V 20/40 - ScenesScene-specific elements in video content
G06V 10/77 - Processing image or video features in feature spacesArrangements for image or video recognition or understanding using pattern recognition or machine learning using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]Blind source separation
G06V 10/80 - Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
75.
TRAINING FOR A MULTIMODAL SPEECH LANGUAGE LARGE MODEL
A training method for a multimodal speech language large model is provided. The implementation is: obtaining first response speech data generated by the multimodal speech language large model by inputting first inquiry speech data into the multimodal speech language large model; determining an inquiry text corresponding to the first inquiry speech data and a response text corresponding to the first response speech data; determining, based on the inquiry text and the response text, a first score; determining, based on speech features of the first inquiry speech data and speech features of the first response speech data, a second score, where the speech features include at least one of speech clarity, speech rate feature, timbre feature, intonation feature, and emotion feature; and adjusting, based on the first score and the second score, parameters of the multimodal speech language large model.
A method for processing a code based on a large model in a human-machine interaction system is performed by an electronic device and includes: obtaining a first dialog text for describing a code processing task; determining at least one of a newly added code file or an original code file as a target code file corresponding to the code processing task; inputting the first dialog text into a large model, and outputting, via the large model, a code processing result of the target code file and a second dialog text corresponding to the first dialog text; and displaying the code processing result and a target dialog text on an interaction interface, wherein the target dialog text comprises at least one of the first dialog text or the second dialog text.
The task execution method includes: retrieving, from a storage unit, a hyperparameter of a target network layer in a target model; executing, using an operator unit, a first computational subtask in a computational task, according to the hyperparameter of the target network layer, so as to obtain a first feature output by the target network layer; executing, in response to reusing the hyperparameter of the target network layer, a second computational subtask in the computational task using the operator unit based on the first feature retrieved from the storage unit, so as to obtain a second feature output by the target network layer, where the first computational subtask and the second computational subtask are subtasks sequentially executed in the computational task; and determining a model output result of the target model using the operator unit based on the second feature retrieved from the storage unit.
Provided is a content generation method and apparatus based on artificial intelligence, a device and a storage medium, relating to the fields of computer vision, deep learning, large model, and intelligent agent. The content generation method includes: sending, by a first intelligent agent, a task execution requirement to a second intelligent agent according to task guidance information, wherein the task guidance information comprises guidance information for generating content, and the task execution requirement comprises a target task that needs to be executed by the second intelligent agent to generate content; and receiving, by the first intelligent agent, a task execution result from the second intelligent agent, wherein the task execution result comprises an execution result generated after the second intelligent agent executes the target task.
Provided are a field-programmable gate array structure for implementing out-of-band bridging, an out-of-band bridging system and method, and a server, relating to the field of computer technology, especially a server. The field-programmable gate array structure includes at least two first data interface modules, at least two second data interface modules, and a buffer module. Each first data interface module is matched to a respective in-band data interface module. Each second data interface module is matched to a respective management data interface module. The buffer module is configured to buffer data of the first data interface modules and data of the second data interface modules. Each second data interface module is configured to perform data transmission with at least two first data interface modules through the buffer module.
A method and an apparatus for determining an operating frequency of a virtual machine, an electronic device, and a storage medium are provided. The method includes: obtaining a target register from a pass-through list, and acquiring register data corresponding to the target register, where registers in the pass-through list is accessible by the virtual machine; during the operation of the virtual machine, in response to receiving an instruction to read the target register data, passing through the register data corresponding to the target register to the virtual machine, so that the virtual machine obtains the operating frequency of the virtual machine according to the register data.
A method and an apparatus for mounting a file system, a device, and a storage medium are provided. The method includes: detecting whether a mount request obtained from a client includes a target token; in response to determining that the target token is included, querying a target authorization record associated with the target token from candidate authorization records of candidate file systems, where the target authorization record includes a target identifier of the target file system; and feeding back the target identifier to the client, enabling the client to access the target file system using the target identifier.
A method for generating corpus data based on at least one large model is provided, which relate to the field of artificial intelligence technologies, and in particular to the fields of deep learning, large models, and intelligent question answering. The method includes: performing a content generation task by using the at least one large model based on a predetermined requirement condition to obtain a corpus content, where the content generation task includes a plurality of target tasks having dependency relationships, and the plurality of target tasks represent a reasoning process of the at least one large model for a corpus content to be generated; and determining target corpus data based on the corpus content and a reasoning process information related to the plurality of target tasks.
A large model-based information processing method, an apparatus, a device, and a medium are provided, which relate to the technical field of artificial intelligence, particularly to the technical fields of machine learning, deep learning, large models and the like. The method includes: obtaining a user input; determining a target working mode from a plurality of predefined working modes, where each predefined working mode has a corresponding inference strategy and is provided with a mode control identifier for triggering the inference strategy; and inputting the user input and the mode control identifier of the target working mode into the large model to obtain target output data generated by the large model based on the inference strategy of the target working mode.
A content sharing method includes: obtaining a first multimedia content to be shared in a terminal during a session between the terminal and a large model; performing a content desensitization processing on a sensitive content in the first multimedia content, and obtaining a second multimedia content by performing a type labeling processing on a position of the sensitive content in the first multimedia content, in which a type labeled by the type label processing is a type of the sensitive content; and sharing with the large model the second multimedia content for the session.
A method for video interaction, an electronic device, and a storage medium are provided. The method may include: during a video interaction with a large model, determining a target object targeted by a spatially directional action associated with a video frame in an interaction process; determining a data processing instruction for the target object based on input information linked to the spatially directional action; and using the large model to perform data processing on the target object according to the data processing instruction, thereby obtaining a data processing result.
G06V 20/40 - ScenesScene-specific elements in video content
G06V 10/80 - Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
G06V 10/86 - Arrangements for image or video recognition or understanding using pattern recognition or machine learning using syntactic or structural representations of the image or video pattern, e.g. symbolic string recognitionArrangements for image or video recognition or understanding using pattern recognition or machine learning using graph matching
G06V 10/94 - Hardware or software architectures specially adapted for image or video understanding
G06V 20/62 - Text, e.g. of license plates, overlay texts or captions on TV images
G06V 20/70 - Labelling scene content, e.g. deriving syntactic or semantic representations
G06V 30/262 - Techniques for post-processing, e.g. correcting the recognition result using context analysis, e.g. lexical, syntactic or semantic context
A model-based peptide design method in the field of artificial intelligence technology such as biological computing is provided. The specific implementation includes: obtaining a pocket of an objective target protein and an objective peptide, a reserved position for designing a unnatural amino acid is identified in the objective peptide, and the pocket binds to the objective peptide via the reserved position; obtaining a feature of the pocket of the objective target protein and multimodal features of each known amino acid in the objective peptide; the multimodal features of each known amino acid comprise a backbone orientation feature, a backbone rotation feature, a side chain type feature, and a rigid atom group distribution feature; designing the unnatural amino acid at the reserved position in the objective peptide using a pre-trained peptide design model based on the feature of the pocket of the objective target protein and the multimodal features of each known amino acid in the objective peptide.
Large model-based visual content generation and target large model training methods, relating to artificial intelligence fields such as deep learning, a large model, computer vision and natural language processing, are provided. A large model-based visual content generation method may include: obtaining target instruction information; inputting the target instruction information into a target large model to obtain and output corresponding target result information, where the target result information includes target visual content, the target result information is generated by the target large model according to target thinking information, and the target thinking information is thinking process information generated by the target large model for the target instruction information.
A method includes: obtaining audio data and a first target image including the face of a target object; performing a facial landmark extraction on the first target image to obtain a first facial landmark image; performing, based on the audio data, an audio feature extraction to obtain an audio feature; inputting the first facial landmark image and the audio feature into a predefined landmark generation network model to obtain a facial landmark image sequence corresponding to the audio data; obtaining, based on the facial landmark image sequence and the first target image, a video corresponding to the audio data generated based on the first target image.
Provided are a method for generating a test case, an electronic device and a storage medium, relating to the field of data processing technology, and in particular to the fields of artificial intelligence, large model and other technologies. The method includes: obtaining test requirement information of code to be tested; calling the large model based on the test requirement information of the code to be tested to generate a target generator for the code to be tested; and obtaining a plurality of target test cases for testing the code to be tested based on the target generator.
A model fusion method includes: calling a main process during a pre-training process of a large model, to cache intermediate model parameters obtained during the pre-training process into a main buffer; and calling a sub-process via the main process to read the intermediate model parameters from the main buffer and perform a parameter fusion process based on the intermediate model parameters.
A method for generating corpus data based on large models is provided, which relates to the field of artificial intelligence technologies, and in particular to the fields of deep learning, large models, and intelligent question answering. The method includes: conducting a dialogue on a predetermined topic by using a plurality of role-based large models to obtain an utterance content of at least one of the role-based large models; performing dialogue strategy planning based on the utterance content by using a designated large model to obtain a dialogue strategy, where the dialogue strategy constrains a speaking pattern of the role-based large models during a dialogue process; and determining target corpus data related to the predetermined topic according to a target utterance content, where the target utterance content is generated by the role-based large models conducting a dialogue based on the dialogue strategy.
A method for generating a live streaming script, an electronic device and a storage medium are provided, which relate to the field of artificial intelligence technologies, in particular to the fields of natural language processing, large models, and virtual digital characters. The method for generating a live streaming script includes: generating at least one first script segment according to an initial input information, where the first script segment includes a speech content text and an object description sub-segment for a live streaming object, the object description sub-segment includes an object description text for describing at least one of an action presented by the live streaming object or a presentation mode for the speech content text; and determining the live streaming script according to the at least one first script segment.
A method for information display based on a large model, a device, and a medium are provided. The method includes: in response to receiving a query question, performing matching between the query question and a field value set of a target data table corresponding to the query question to obtain at least one target field value group, where the target field value group includes target field values semantically associated with each other, and the target field values are configured to reduce a semantic deviation in the large model's understanding of the query question; invoking the large model according to a prompt information to generate a target query statement, where the prompt information is obtained based on the query question, the at least one target field value group, and a description information of the target data table; and displaying a query result obtained by executing the target query statement.
Provided is a task-oriented dialogue implementation method relating to artificial intelligence fields such as deep learning, large language models, natural language processing and intelligent agents, which can be applied to intelligent interaction scenarios such as intelligent customer service, intelligent outbound calling, and intelligent marketing. The task-oriented dialogue implementation method may include: acquiring a question to be answered; generate an answer corresponding to the question by using a task-oriented dialogue model, the answer is generated by the task-oriented dialogue model according to model configuration information, the model configuration information includes a dialogue flow corresponding to the task-oriented dialogue model, and the dialogue flow is written in text form.
A method for generating a digital human video based on a large model, an electronic device, and a storage medium are provided, which relate to a field of artificial intelligence technologies, and may be applied to scenarios such as video livestreaming, advertisement production, and e-commerce sales. The method includes: acquiring a requirement information including an action description information for describing a specified action video segment, and the action video segment represents a specified action of a target object; processing the requirement information using a first large model to obtain a target script, where the target script includes a target speech segment text matching the action description information; and processing the target script and the action video segment using a second large model to obtain a target video for displaying a target digital human performing a speech delivery based on the target speech segment text while performing the specified action.
The present disclosure provides a system and a method for vehicle sensor time synchronization, and a field programmable gate array chip, in the field of artificial intelligence, such as sensor, artificial intelligence chip, and autonomous driving. The system for vehicle sensor time synchronization includes: target vehicle sensors and a field programmable gate array chip, the target vehicle sensors include all vehicle sensors in a target vehicle that require a timestamp information to be written; the target vehicle sensors are connected to the field programmable gate array chip, and are configured to send collected sensor data to the field programmable gate array chip; the field programmable gate array chip is configured to write the timestamp information into the sensor data based on a system time of the field programmable gate array chip; wherein the field programmable gate array chip comprises: a system module and a logic module; the system module and the logic module are respectively connected to at least one of the target vehicle sensors based on a principle of load balancing, and are respectively configured to write the timestamp information into obtained sensor data based on the system time.
Provided is a tensor processing method, an electronic device, and a storage medium, relating to the fields of deep learning and artificial intelligence. The method includes: determining relevant information of a conversion function corresponding to each of one or more target input tensors of a first operator in a target computation graph based on computation logic of the first operator and source split states of at least part of source input tensors of the first operator; splitting each source input tensor of the first operator based on the relevant information of the conversion function corresponding to each target input tensor to obtain each target input tensor; and sending each target input tensor to a plurality of computing devices. The plurality of computing devices are configured to perform distributed parallel communication based on each target input tensor and the first operator, to obtain an output tensor of the first operator.
A method for analyzing a ball game motion includes: in response to entering a serving stage, obtaining a motion trajectory of a ball by fitting based on a ball game video acquired; determining a hitting coordinate of a motion subject, a hitting posture of the motion subject, and a landing coordinate of the ball based on the motion trajectory and the ball game video; and obtaining a motion analysis result based on the hitting coordinate, the hitting posture, and the landing coordinate, to display the motion analysis result in real time through a display device.
A method for training a text question and answer (Q&A) model is performed by an electronic device. The method includes: determining a sample question text set and a sample answer text corresponding to a sample question text in the sample question text set; inputting the sample question text into a text Q&A model to be trained, and obtaining a predicted answer text output by the text Q&A model and at least one prediction probability of at least one reference character on each character position in the predicted answer text; determining an uncertainty degree of the predicted answer text; and obtaining a trained text Q&A model by adjusting a parameter of the text Q&A model based on the sample answer text, the predicted answer text and the uncertainty degree of the predicted answer text.
Provided is a quantization parameter storage method, a model inference method, an electronic device and a storage medium, relating to the fields of large model technology, artificial intelligence technology and model quantization technology. The quantization parameter storage method includes: obtaining, by a calculation unit of a processor, a statistical value of a first quantization parameter of a model statistically based on benchmark data; searching for, by the calculation unit, a target value of the first quantization parameter and a target value of a second quantization parameter of the model in a search space based on the statistical value of the first quantization parameter; and storing, by the calculation unit, the target value of the first quantization parameter and the target value of the second quantization parameter into a memory.