Skip to main content

Connecting SharePoint to Knowledge Management

The SharePoint knowledge integration syncs content from your SharePoint Online sites into the Rezolve.ai Knowledge Base. The Virtual Agent can then use that content to answer questions and cite the source document. Agents can also use it as a reference.

The integration connects through the Microsoft Graph API using an Azure (Entra ID) app registration that you own. You can sync:

ContentWhat gets ingested
Drives (document libraries)Every supported file in the library, including all sub-folders
FoldersEvery supported file in the folder and its sub-folders
Site PagesThe text of modern SharePoint pages (SitePages/*.aspx), either all pages of a site or selected pages
Sub-sitesContent of the selected sub-sites, either the whole sub-site or specific libraries, folders, pages, and lists within it
ListsEach list item becomes a knowledge entry. You choose which columns are searchable.

This page covers the SharePoint-specific setup. For the wizard, sync status, bounced files, and other functions that all integrations share, see Knowledge Integrations.


Prerequisites​

  • SharePoint Online (Microsoft 365 commercial cloud). SharePoint Server (on-premises) and the national clouds (GCC High, DoD, China/21Vianet) are not supported.
  • An Azure / Entra ID administrator who can register an application and grant admin consent.
  • For Sites.Selected only: a SharePoint or Global administrator to grant the app access to each site.
  • Knowledge Management admin access in the Rezolve.ai console.

Step 1: Register an Azure application​

  1. In the Azure portal, open App registrations and click + New registration.
  2. Enter a name, for example Rezolve.ai SharePoint Ingestion. Keep the default single-tenant account type, then click Register.
  3. On the Overview page, copy these two values:
    • Directory (tenant) ID, which is the Azure Tenant ID in Rezolve.ai
    • Application (client) ID, which is the Client ID in Rezolve.ai
  4. Open Certificates & secrets → Client secrets → + New client secret. Enter a description, choose an expiry, and click Add.
  5. Copy the secret Value right away. Azure shows it only once.
note

Only client secret authentication is supported. Certificate-based authentication is not. Note the secret's expiry date: when it expires, syncs fail until you update the connection with a new secret.

For a screenshot walkthrough, see How to Register an Azure Application for SharePoint Ingestion.


Step 2: Grant Microsoft Graph permissions​

Rezolve.ai only reads from SharePoint. Choose one of the two permission models below. Both are Application permissions (not Delegated) and both need admin consent.

Permission modelAccess grantedWhen to use
Sites.Read.AllRead access to every site in the tenantSimplest setup, or when you'll sync many sites
Sites.Selected (recommended)Read access only to sites you explicitly grantLeast-privilege setup. Access is limited to the sites you sync.
info

The Rezolve.ai app doesn't need the Files.Read.All, Sites.ReadWrite.All, or Sites.FullControl.All Graph permissions. Don't grant them to the ingestion app. The one exception is a site-level fullcontrol grant under Sites.Selected, which is only needed for permission inheritance.

Option A: Sites.Read.All​

  1. In your app registration, open API permissions → + Add a permission → Microsoft Graph → Application permissions.
  2. Search for and select Sites.Read.All, then click Add permissions.
  3. Click Grant admin consent for <your directory>.

That's all. The app can now read every site.

Option B: Sites.Selected (per-site access)​

With Sites.Selected, consent alone gives the app no access. You must also grant it read access on each site you want to sync.

  1. In your app registration, add the Microsoft Graph → Application → Sites.Selected permission and click Grant admin consent.

  2. Find the ID of each site. For example, in Graph Explorer:

    GET https://graph.microsoft.com/v1.0/sites/yourcompany.sharepoint.com:/sites/HR

    The id in the response looks like yourcompany.sharepoint.com,<guid>,<guid>.

  3. Grant the Rezolve.ai app read access to the site. Run this as an administrator with Sites.FullControl.All, using Graph Explorer, PowerShell, or a separate admin app. Do not use the Rezolve.ai ingestion app for this step.

    curl --location 'https://graph.microsoft.com/v1.0/sites/<SITE_ID>/permissions' \
    --header 'Authorization: Bearer <ADMIN_ACCESS_TOKEN>' \
    --header 'Content-Type: application/json' \
    --data '{
    "roles": ["read"],
    "grantedToIdentities": [
    {
    "application": {
    "id": "<CLIENT_ID>",
    "displayName": "<APP_DISPLAY_NAME>"
    }
    }
    ]
    }'
    PlaceholderValue
    <SITE_ID>The site ID from step 2
    <ADMIN_ACCESS_TOKEN>A token for an admin identity that holds Sites.FullControl.All
    <CLIENT_ID>The Application (client) ID of the Rezolve.ai ingestion app
    <APP_DISPLAY_NAME>The display name of the Rezolve.ai ingestion app
  4. Repeat step 3 for every site collection you want to sync.

    • Classic sub-sites under the site's URL, such as /sites/HR/Benefits, are covered by the grant on the parent site.
    • Sites that are separate site collections need their own grant, even when they appear related, for example hub-associated sites such as /sites/HR-Policies.
Using Permission Inheritance with Sites.Selected

Permission inheritance reads each file's sharing permissions. SharePoint returns the complete list only to site owners, so grant the fullcontrol role instead of read on sites where you enable Inherit Permissions from SharePoint. The write role is not enough. The app still only reads content.

caution

Don't add Sites.Read.All or Files.Read.All to an app that uses Sites.Selected. Those permissions give access to every site and override the per-site restriction.

To grant access with PnP PowerShell instead of the Graph API, see Sites.Selected - Enabling API Permissions for SharePoint Sites Crawl with PowerShell Script.

caution

If the app can't read a site, library, or page, Rezolve.ai skips it without reporting an error. When a sync finishes but content is missing, check the permission grant first. See Troubleshooting.


Step 3: Create the SharePoint integration​

  1. In the Rezolve.ai console, go to Knowledge Management → Knowledge Integrations.
  2. Click Add Integrations and select SharePoint.

3a. Connection setup​

Enter a Title for the integration. Then either choose an existing SharePoint connection (Choose From Existing Connections) or select Create New Connection and enter these details:

FieldDescriptionExample
NameA descriptive name for the connectionContoso SharePoint - HR
Site URLFull URL of the SharePoint site to synchttps://yourcompany.sharepoint.com/sites/HR
Azure IdDirectory (tenant) ID from Step 172f988bf-86f1-41af-91ab-2d7cd011db47
Client IdApplication (client) ID from Step 1bf7f3a64-7c3c-4c71-91a2-2b50fd0a7f2c
Client SecretClient secret value from Step 1 (not the secret ID)~2TjQ8...

Site URL rules:

  • Use the full URL of a site under /sites/ or /teams/. Any other URL form is treated as the tenant root site.
  • Use the *.sharepoint.com address. Custom or vanity domains are not resolved.
  • The Site URL can't be edited after the integration is created. To sync a different site, create a new integration.

When editing a connection later, you can replace the Client Id and Client Secret, for example when the secret expires.

3b. Content selection​

Content is organized into one or more content sets. Each content set has its own selection and access settings. In each content set, browse the site tree and tick what to sync. You don't need to type any IDs.

Item in the treeWhat it syncs
Drives (document libraries)The whole library, crawled recursively. Expand a drive to pick individual folders instead. The Preservation Hold Library is never synced.
FoldersThe folder and all of its sub-folders. A folder inside a ticked drive is already covered and isn't processed twice.
PagesTick the Pages folder to sync all site pages, including pages published later. Expand it to pick individual pages; pages load 50 at a time, and Load More shows the rest. Individually picked pages stay a fixed list. Unticking any page switches the content set to an individual page list.
Sub-sitesTick a sub-site to include its content, or expand it to pick its drives, folders, pages, lists, and nested sub-sites individually.
ListsShown when the SharePoint List switch in the content set's ⋮ menu is on. See SharePoint lists.

Restricting access​

Open the content set's ⋮ menu and turn on Restrict Access. Then choose either or both of:

  • Restrict by Audience: pick one or more Audience(s) that can get answers from this content set.
  • Inherit permissions from SharePoint: restrict each item to the SharePoint users who can access it. See Permission inheritance.

Without Restrict Access, content is available to everyone who uses the Virtual Agent.

SharePoint lists​

To sync SharePoint lists, open the content set's ⋮ menu and turn on SharePoint List. This switch appears for users with the Knowledge admin role. The site's lists then appear in the tree. Only visible custom lists that have visible columns are offered.

Tick a list and expand it to choose its fields:

SettingPurpose
Contextual FieldsThe columns whose values become the searchable text of each item. Select at least one.
Meta FieldsColumns stored as metadata alongside each item. Remove any you don't need.

Each list row becomes one knowledge entry that links to the item's display form (DispForm.aspx?ID=...).

3c. Configuration​

Under Sync Details, set the Frequency (daily, weekly, or monthly), Start Time, Timezone, and Schedule Start Date. Optionally turn on Subscribe to Sync Reports and add User Subscribers. Then click Save. See Knowledge Integrations → Step 3: Configuration.


What gets ingested​

Supported file types​

CategoryExtensions
Documents.pdf, .doc, .docx, .txt, .md
Presentations.ppt, .pptx
Email.msg
Archives.zip, .rar, .7z (supported files inside are extracted)
PagesModern site pages (.aspx)

Excel (.xls, .xlsx), CSV, images, video, and audio files in document libraries are not ingested. They appear as bounced files with an unsupported file type reason.

Limits​

LimitDefault
Maximum file size25 MB per file
Archive contentsUp to 10 files and 50 MB uncompressed per archive
Retries per fileA file that fails 3 times, or fails for a reason that won't change, is skipped on later syncs until it's resolved in Bounced Files

Site page content​

  • Only the text web parts of a page are ingested. Images, embedded videos, lists, and other non-text web parts are ignored.
  • Links in page text are kept and made absolute.
  • A page with no text is skipped. It isn't treated as a failure.

Metadata​

Each ingested document keeps:

  • its name
  • its SharePoint URL, which the Virtual Agent shows as the source link
  • its last-modified time
  • its SharePoint column values, including columns inherited from parent folders
note

Draft or checked-out versions aren't filtered out. Whatever the app can read in the library is ingested. Keep drafts in a location you don't sync.


Sync behavior​

BehaviorDetails
First syncFull ingestion of everything selected
Scheduled / Sync NowIncremental. Each run compares the selected locations with what's already ingested, then adds new items, re-ingests items modified since the last sync, and removes items that no longer exist in SharePoint.
Moves and renamesTreated as a delete at the old location plus an add at the new one
Removing a source from a content setWhen you save, the content from the deselected source is removed from the Knowledge Base. If you deselect a drive or sub-site, a sync also starts automatically.
Other content changesNewly added sources are synced on the next scheduled sync, or right away with Sync Now
Stop SyncStops future syncs. Content already ingested stays in the Knowledge Base.
Delete integrationRemoves the configuration and all content it ingested
ThrottlingWhen Microsoft Graph throttles requests (HTTP 429) or has temporary errors, requests are retried automatically, respecting the Retry-After header

Each run enumerates every selected location, so very large libraries take longer to sync even when little has changed. Schedule large syncs outside business hours.


Permission inheritance​

When Inherit Permissions from SharePoint is enabled, the permissions below are recorded with each item. The Virtual Agent then only returns content to users who have access to it.

ContentPermissions recorded
Files in drives and foldersThe SharePoint groups and Entra ID groups that have read, edit, or owner access to the file
Site pagesSite-level access only. Per-page unique permissions are not captured.
List itemsThe list's configured audience column, if one is set
caution

Access granted directly to individual users, rather than through a group, isn't captured. For accurate access control, manage SharePoint permissions through groups.


Troubleshooting​

SymptomLikely causeResolution
Sync succeeds but a site, library, or page is missingThe app has no access to it. This is common with Sites.Selected when a separate site collection (for example a hub-associated site) wasn't granted.Grant read access to that site (Step 2, Option B), then click Sync Now
All syncs fail after working earlierThe client secret has expiredCreate a new secret in Azure and update the connection's Client Secret
Authentication error when saving the connectionWrong tenant ID, client ID, or secret. The secret ID was entered instead of the secret value.Re-copy the values from the app registration
Content from the tenant root site appears instead of the expected siteThe site URL isn't in /sites/... or /teams/... formUse the full https://<tenant>.sharepoint.com/sites/<name> URL
Site URL not foundA custom or vanity domain was usedUse the *.sharepoint.com URL
File listed in Bounced Files as size limit exceededFile is over 25 MBSplit or compress the file, or exclude it
File listed as unsupported file typeFor example, Excel, CSV, or image filesConvert to a supported format (for example, PDF) if the content is needed
Permission inheritance is on but users can't get answers they should see, or the audience looks incompleteWith Sites.Selected, the app has only the read role, so file permissions come back incompleteChange the site grant to fullcontrol, then click Sync Now
Scanned PDF ingested with little or no textText in images isn't readUse a PDF with a text layer, for example by running OCR in Acrobat before uploading
Page ingested with missing contentThe content is in image, embed, or other non-text web partsPut the key information in text web parts
A file stopped retryingIt failed 3 timesFix the file in SharePoint, resolve it in Bounced Files, then click Sync Now

Download the CSV log for any sync run from the Sync Status tab. See Downloading Sync Logs.


Resources​

ResourceLink
Register an application in Entra IDMicrosoft Learn
Graph permissions reference (Sites)Microsoft Learn
Selected permissions in SharePoint (Sites.Selected)Microsoft Learn
Create site permission (Graph API)Microsoft Learn
Microsoft Graph SharePoint APIMicrosoft Learn

Next Steps​