# Faster bulk\_create using dictionaries

**URL:** <https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891>\
**Category:** ORM\
**Created:** [January 10, 2026, 3:49pm UTC](https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891 "2026-01-10T15:49:05Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![adamsol](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.djangoproject.com/adamsol/32/31184_2.png) [@adamsol](https://forum.djangoproject.com/u/adamsol)\
**Post date:** [January 10, 2026, 3:49pm UTC](https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891/1 "2026-01-10T15:49:05Z")

</div>

Recently, there were some optimizations in `bulk_create`:

- [#35936 (Speeding up Postgres bulk\_create by using unnest) – Django](https://code.djangoproject.com/ticket/35936)
- [#36088 (Avoid unnecessary DEFAULT usage on bulk\_create for models with db\_default fields) – Django](https://code.djangoproject.com/ticket/36088)
- [#36815 (Avoid unnecessary prepare\_value calls when inserting db\_defaults) – Django](https://code.djangoproject.com/ticket/36815)

However, one of the slowest parts of the whole process of bulk-inserting data is still (quite surprisingly) creating Python objects before the actual inserting. This was already discussed in other threads:

- [How to avoid the overhead of model instances in bulk\_create](https://forum.djangoproject.com/t/how-to-avoid-the-overhead-of-model-instances-in-bulk-create/25538)
- [Speeding up Postgres bulk\_create by using unnest - #12 by jerch](https://forum.djangoproject.com/t/speeding-up-postgres-bulk-create-by-using-unnest/36508/12)

The same issue also significantly [slows down adding identifiers to M2M fields](https://code.djangoproject.com/ticket/31202#comment:16).

(For additional reference, see [this SO thread](https://stackoverflow.com/a/55256047) for a benchmark of creating objects in Python, or [this blog post](https://johnnymetz.com/posts/django-values-over-only/) for a report of huge performance gains by avoiding creating model instances when fetching lots of data.)

So I was thinking about an option of inserting data with dictionaries. Since we already have `values` for fetching rows using dictionaries instead of objects, inserting data this way should fit as well. It could be handled directly in the existing `bulk_create` method (allowing for an arbitrary mix of dictionaries and objects in the list). We can just take data from the dictionary instead of from the object, and for missing keys we can insert the model defaults.

I’ve been experimenting with this, and have prepared an example implementation as a proof of concept: [Faster bulk\_create using dictionaries · adamsol/django@ed1ad9c · GitHub](https://github.com/adamsol/django/commit/ed1ad9c1e53a6ac23df2474f71bd997401dfc93c). The speedup in my tests was between 1.6x and 2x, depending on the model. It seems that models with `db_default` benefit the most, as we additionally avoid creating `DatabaseDefault` instances for each object.

Does this sound like something worth implementing in Django?

It’s of course possible to build such a helper function outside of Django (which I have done in a project I’m working on). Nonetheless, it would be convenient to have a faster method of inserting data built into the framework - especially for DB migrations, as they tend to be difficult to test automatically, so importing and using custom functions may easily lead to them getting broken after some code refactoring.

---

<div class="post-metadata">

**Author:** ![jerch](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.djangoproject.com/jerch/32/13571_2.png) [@jerch](https://forum.djangoproject.com/u/jerch)\
**Post date:** [January 11, 2026, 4:35pm UTC](https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891/2 "2026-01-11T16:35:38Z")

</div>

@adamsol I have not looked yet on your approach, but want to give you a few more pointer on what i have tried so far and why.

For django-computedfields I tested different approaches to cut ORM functionality which these results:

> <https://github.com/netzkolchose/django-computedfields/issues/195>
>
> This is just a playground for now...
> 
> The limitations around UNIONed queries and… the .only idea made me think, whether we can safe update runtime by further cutting ORM ropes.
> 
> My first idea was to leave out the instance creation in \_bulk\_updater\_ and just go with dictionaries for local updates. So I took the most heavy local cf model \_SelfRef\_ and edited \_bulk\_updater\_ and the compute functions to refer to dict entries instead of instance attributes with these results in \_updatedata\_ command (for 100k records):
> \- run with no update: 48000 rec/s --\> 110000 rec/s, speedup is ~2.2x
> \- run with full update: 20000 rec/s --\> 30000 rec/s, speedup is ~1.5x
> 
> The benefit is somewhat underwhelming given the fact, that it drops all nice attribute access pattern the ORM provides. Or to put it differently - while the ORM instance creation puts a significant perf burden in select queries, it is still only within a ~2x range. So the ORM does a pretty good job not to penalize things too much.
> 
> Well, next stop for investigations is the raw cursor interface (the tests above were still using the queryset.values() ORM interface) ...

This tested SELECTs with model instances vs. dictionaries retrieved via `.values()`. The roundtrip with updates would still create model instances, thus the benefit is lower with ~1.5 times speedup. For a fully dictionary-based handling I expect the benefit somewhere in 2-3x speedup for postgres. This at least is indicated by my tests with `copy_insert` impls tested here: [idea - should the postgres copy path get a copy\_insert/create method? · Issue #4 · netzkolchose/django-fast-update · GitHub](https://github.com/netzkolchose/django-fast-update/issues/4)

I had no time yet to write everything down into neatly tested lib code yet, as I got distracted with a few psycopg issues and patched those first.

---

<div class="post-metadata">

**Author:** ![adamchainz](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.djangoproject.com/adamchainz/32/26_2.png) [@adamchainz](https://forum.djangoproject.com/u/adamchainz)\
**Post date:** [January 11, 2026, 11:32pm UTC](https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891/3 "2026-01-11T23:32:31Z")

</div>

> [@adamsol](#):
>
> Does this sound like something worth implementing in Django?

Maybe! I agree it fits with how `.values()` can return dictionaries—the symmetry of allowing `dict`s in some operations is appealing. However, it would be a big scope change, since it would logically lead to`bulk_update` also accepting dictionaries.

I would also like to see attempts to optimize `Model. __init__ ` so this is less of a problem. It does a lot of work, and while some improvements have been made, and maybe there are more yet.

> [@adamsol](#):
>
> It seems that models with `db_default` benefit the most, as we additionally avoid creating `DatabaseDefault` instances for each object.

Aha, that’s a good hint… I found that we don’t need to create `DatabaseDefault` instances per model instance, leading to this ~12% optimization: [#36858 (Optimize `db_default` creation) – Django](https://code.djangoproject.com/ticket/36858)

---

<div class="post-metadata">

**Author:** ![adamsol](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.djangoproject.com/adamsol/32/31184_2.png) [@adamsol](https://forum.djangoproject.com/u/adamsol)\
**Post date:** [January 12, 2026, 6:45pm UTC](https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891/4 "2026-01-12T18:45:53Z")

</div>

> I would also like to see attempts to optimize `Model. __init__ ` so this is less of a problem. It does a lot of work, and while some improvements have been made, and maybe there are more yet.

I’m not sure if much can be done on the Django side here, since object creation overhead comes from Python itself - as benchmarked in the [SO answer](https://stackoverflow.com/a/55256047) that I linked earlier. It’s getting better in newer Python versions (my measurements were on 3.13), but dictionaries should still win convincingly in most cases.

> I found that we don’t need to create `DatabaseDefault` instances per model instance, leading to this ~12% optimization:

Nice, so now the advantage of avoiding objects will diminish a little, but 1.6x-1.8x should still be achievable.

---

<div class="post-metadata">

**Author:** ![adamchainz](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.djangoproject.com/adamchainz/32/26_2.png) [@adamchainz](https://forum.djangoproject.com/u/adamchainz)\
**Post date:** [January 12, 2026, 8:48pm UTC](https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891/5 "2026-01-12T20:48:56Z")

</div>

> [@adamsol](#):
>
> as benchmarked in the [SO answer](https://stackoverflow.com/a/55256047) that I linked earlier.

Yeah, objects cannot be as fast as plain dicts. But the thread is not quite an apples-to-apples comparison, as it’s using dataclasses, which have generated code that may not be as efficient as a vanilla class can be.

> [@adamsol](#):
>
> since object creation overhead comes from Python itself

`Model. __init__ ` does a _lot_ of stuff: [django/django/db/models/base.py at 2b192bff26cf956c168790fce6a637cbd814250b · django/django · GitHub](https://github.com/django/django/blob/2b192bff26cf956c168790fce6a637cbd814250b/django/db/models/base.py#L502)

I optimized it a bit ten years ago in [Optimized Model instantiation a bit. · django/django@d2a26c1 · GitHub](https://github.com/django/django/commit/d2a26c1a90e837777dabdf3d67ceec4d2a70fb86)

I’m sure some more targeted profiling and investigating could find further speedups, especially given how Django and Python have changed since then.

(This is not to dismiss the `dict` support idea, still.)

---

<div class="post-metadata">

**Author:** ![adamsol](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.djangoproject.com/adamsol/32/31184_2.png) [@adamsol](https://forum.djangoproject.com/u/adamsol)\
**Post date:** [January 13, 2026, 6:05pm UTC](https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891/6 "2026-01-13T18:05:21Z")

</div>

You’re right, after benchmarking further, I can see that `Model. __init__ ` is actually the main culprit, and Python’s object creation overhead is less significant. The following script:

```py
def measure(f):
    import time
    t = time.perf_counter()
    f()
    print(f'{(time.perf_counter() - t):.3f}')

class A:
    def __init__ (self, id):
        self.id = id

class B(models.Model):
    class Meta:
        app_label = 'test'

N = 100_000

measure(lambda: [{'id': i} for i in range(N)])
measure(lambda: [A(id=i) for i in range(N)])
measure(lambda: [B(id=i) for i in range(N)])

```

gives results like:

```auto
0.018
0.051
0.236

```

But I guess this doesn’t change much regarding the dictionary idea.

---

<div class="post-metadata">

**Author:** ![adamsol](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.djangoproject.com/adamsol/32/31184_2.png) [@adamsol](https://forum.djangoproject.com/u/adamsol)\
**Post date:** [January 24, 2026, 1:04pm UTC](https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891/7 "2026-01-24T13:04:45Z")

</div>

I’ve created [Faster bulk\_create using dictionaries · Issue #113 · django/new-features · GitHub](https://github.com/django/new-features/issues/113)

---

<div class="post-metadata">

**Author:** ![jerch](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.djangoproject.com/jerch/32/13571_2.png) [@jerch](https://forum.djangoproject.com/u/jerch)\
**Post date:** [March 20, 2026, 5:59pm UTC](https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891/8 "2026-03-20T17:59:27Z")

</div>

@adamsol Had a quick look at your demo impl - looks pretty straight forward, nice.

Still I stumbled over the value adaption your are doing in your approach as:

```python
# your direct approach
value = obj[field.attname]
# vs. django's adaption
value = field_pre_save(obj)

```

I think this will lead to different behavior for certain field types and values, as some field types do advanced value adaption in the pre\_save logic (e.g. JSON-field’s None behavior). Furthermore the docs mentions this as the way to do value adaptions.

While I like the idea to skip individual field adaption (creates a huge amount of runtime for postgres DB), I think this needs to be discussed, whether it should take that route (then plus docs with a hint, that ppl have to adapt values prehand in their dicts) or whether some basic adaption to flat out DB differences should be still done here.

---

<div class="post-metadata">

**Author:** ![adamsol](https://sea2.discourse-cdn.com/flex026/user_avatar/forum.djangoproject.com/adamsol/32/31184_2.png) [@adamsol](https://forum.djangoproject.com/u/adamsol)\
**Post date:** [March 21, 2026, 11:34am UTC](https://forum.djangoproject.com/t/faster-bulk-create-using-dictionaries/43891/9 "2026-03-21T11:34:43Z")

</div>

If I’m seeing correctly, all adaptations actually happen in `field_prepare` (`prepare_value`), which my code still calls. My code indeed doesn’t call `field_pre_save` (`pre_save_val`), but that function is only for things like `auto_now` and files. I think the reason I skipped it was because `pre_save_val` is intertwined with calling `getattr` on the object, which cannot work in the context of dictionaries. Also, something similar already happens for raw queries: there is a condition in `pre_save_val` to avoid calling `pre_save` for them. So generally some decision would be required here: whether `bulk_create` via dictionaries is supposed to be more like a raw query or a standard query. And either some more changes in the code would be necessary, or the difference would need to be documented.
